📊 Medical Imaging AI Weekly Papers

Segmentation · Generation · Detection · AI Agent · Registration · Dose Calculation
🕐 2026-08-13 08:12:01 (UTC+8)
📅 2026-08-13 📄 32 Papers 🤖 arXiv API + AI Summary
📚 Archive

🔬 Medical Image Segmentation

5条
Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framew
👤 Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd et al. (11 authors) 📅 2026-08-07 🔗 arXiv 📄 PDF
The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors from large volumes to characterize pathologies. Da
👤 Robin Trombetta, Carole Lartizien 📅 2026-08-06 🔗 arXiv 📄 PDF
Deformable image registration (DIR) is a core problem in medical image analysis; but, unlike labeling decision problems such as classification and segmentation, registration is a problem class that in
👤 Onur Ali Zeybekoglu, David Tilly, Orcun Goksel 📅 2026-08-03 🔗 arXiv 📄 PDF
Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain d
👤 Ke Ma, Yamin Mao, Weiming Li, Shuai Tan et al. (8 authors) 📅 2026-08-11 🔗 arXiv 📄 PDF
Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often model global and local features in separate modules o
👤 Zheyang Jing, Qin Lu, Jianwang Li, Yujie Yang et al. (6 authors) 📅 2026-08-09 🔗 arXiv 📄 PDF

🤖 Medical AI Agent & VLM

23条
Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not
👤 Shengzhi Wang, Jun Yang, Kai Wu, Xiaozhong Ji et al. (14 authors) 📅 2026-08-11 🔗 arXiv 📄 PDF
While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables. However, existing m
👤 Yingsheng Liu, Haiming Li, Jingmin Zhu, Jiajun Sun et al. (9 authors) 📅 2026-08-11 🔗 arXiv 📄 PDF
Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, w
👤 Haoyu Yang, Meixing Shi, Zengjie Chen, Haoran Sun et al. (8 authors) 📅 2026-08-10 🔗 arXiv 📄 PDF
Surgical instrument segmentation is a fundamental task for computer-assisted interventions, yet most existing methods rely on pixel-level annotations or manual spatial prompts, which limit scalability
👤 Nakul Poudel, Richard Simon, Cristian A. Linte 📅 2026-08-09 🔗 arXiv 📄 PDF
Despite significant advances in Medical Report Generation (MRG), the reliability remains constrained by the prevalence of factual errors. While Direct Preference Optimization (DPO) has emerged as a pr
👤 Qiang Hu, Yuxuan Luo, Yingjie Guo, Hao Wang et al. (7 authors) 📅 2026-08-09 🔗 arXiv 📄 PDF
Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter count has been the dominant efficiency metric in PE
👤 Uri Z. Kialy, Gil Ben-Artzi 📅 2026-08-08 🔗 arXiv 📄 PDF
Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical c
👤 Jiaxuan Li, Qing Xu, Xiangjian He, Yue Li et al. (7 authors) 📅 2026-08-06 🔗 arXiv 📄 PDF
Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purp
👤 Zihan Wang, Tong Liu, Zhiwei Wang, Tao Huang et al. (9 authors) 📅 2026-08-04 🔗 arXiv 📄 PDF
Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attention maps provide only local explanations: they indic
👤 Abdallah Lamane, Abdul Rahman Diab, Ren-Chin Wu, William Lotter 📅 2026-08-03 🔗 arXiv 📄 PDF
Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain M
👤 Amir Sabbaghziarani, Mohammadsajad Abavisani, Sergey Plis 📅 2026-08-03 🔗 arXiv 📄 PDF
Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the dominant pretraining source. While effective at s
👤 Haozhe Luo, Ziyu Zhou, Shelley Zixin Shu, Mauricio Reyes 📅 2026-08-03 🔗 arXiv 📄 PDF
Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imaging data and task-specific training. We investigat
👤 Md Rabiul Islam, Samir Abdaljalil, Erchin Serpedin, Hasan Kurban 📅 2026-08-11 🔗 arXiv 📄 PDF
Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-
👤 Ying Jin, Noel C. F. Codella, John Corring, Mu Wei et al. (6 authors) 📅 2026-08-11 🔗 arXiv 📄 PDF
Public chest X-ray repositories are widely used to train medical AI systems, yet their labels are typically extracted from radiology reports rather than verified directly on images. As a result, repos
👤 Yesika Alexandra Agudelo-Londoño, Jhon Wilmer Pino-Román, Brahian Carrera Rodríguez, José Miguel Castañeda-Bedoya et al. (12 authors) 📅 2026-08-10 🔗 arXiv 📄 PDF
Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges. Modern foundation visio
👤 Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer et al. (13 authors) 📅 2026-08-09 🔗 arXiv 📄 PDF
Text-conditioned chest X-ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work has primarily improved quality by updating image generators, i
👤 Pengxiang Cai, Xiaohan Li, Anglin Liu, Qingyuan Zeng et al. (6 authors) 📅 2026-08-08 🔗 arXiv 📄 PDF
Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, prior,uncertain, or irrelevant findings, while multiple valid ab
👤 Chengyi Peng, Haoyu Yang, Meixing Shi, Yuxiang Cai et al. (5 authors) 📅 2026-08-07 🔗 arXiv 📄 PDF
Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for Vision-Language Models (VLMs). We benchmark stat
👤 Bruno Palau, Franziska Vogt, Daria Laslo, Haobo Li et al. (7 authors) 📅 2026-08-07 🔗 arXiv 📄 PDF
Computed tomography (CT) is widely used for clinical diagnosis and longitudinal follow-up, yet automatically generating accurate and complete radiology reports from three-dimensional (3D) CT remains c
👤 Dongchen Li, Jitao Liang, Wei Li 📅 2026-08-06 🔗 arXiv 📄 PDF
Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported
👤 Yuta Kobayashi, Pradyun Ramesh, Muhammad Ahmed Chaudhry, Vincent Jeanselme et al. (8 authors) 📅 2026-08-05 🔗 arXiv 📄 PDF
A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measure
👤 Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi et al. (11 authors) 📅 2026-08-04 🔗 arXiv 📄 PDF
Automatic radiology report generation (RRG) aims to simulate the workflow of radiologists, assisting them in clinical diagnosis. However, existing methods often fall short in utilizing all information
👤 Yang Yu, Yiming Ji, Bin Dai, Dong Zhang et al. (7 authors) 📅 2026-08-04 🔗 arXiv 📄 PDF
In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, using short-form findings from the reports. In this study, we dev
👤 Ruida Cheng, Tejas S. Mathai, Benjamin Hou, Qingqing Zhu et al. (7 authors) 📅 2026-08-03 🔗 arXiv 📄 PDF

🔄 Medical Image Registration

3条
Registration-based Few-shot medical image segmentation (RFMIS) aims to generate pseudo-labels for unlabeled images by warping a labeled image through registration. However, existing methods primarily
👤 Jia Wang, Jiaming Cai, Zunying Hu, Zhanjie Wu et al. (7 authors) 📅 2026-08-07 🔗 arXiv 📄 PDF
Anatomical image registration commonly relies on a sequential pipeline where an affine alignment is estimated first and then held fixed while a non-rigid diffeomorphic deformation is applied. This two
👤 Anton François, Rayane Mouhli, Thomas Pierron 📅 2026-08-11 🔗 arXiv 📄 PDF
Longitudinal MRI enables sensitive measurement of structural brain change for studying aging and neurodegenerative disease. Deformable image registration is a key tool for estimating such change by co
👤 Jingru Fu, Kathleen E. Larson, Douglas N. Greve, Bruce Fischl et al. (5 authors) 📅 2026-08-08 🔗 arXiv 📄 PDF

☢️ Radiation Dose Calculation

1条
During standard radiotherapy planning, repeated CT acquisitions are often required for patient registration, verification, and adaptive planning, resulting in increased cumulative X-ray dose. To mitig
👤 Alzahra Altalib, Chunhui Li, Christopher Hamill Taylor, Sankar Pillai et al. (5 authors) 📅 2026-08-09 🔗 arXiv 📄 PDF