📊 Medical Imaging AI Weekly Papers

Segmentation · Generation · Detection · AI Agent · Registration · Dose Calculation
🕐 2026-07-30 08:11:14 (UTC+8)
📅 2026-07-30 📄 21 Papers 🤖 arXiv API + AI Summary
📚 Archive

🔬 Medical Image Segmentation

3条
Segmentation of adjacent structures with similar intensity distributions remains a challenging problem in image analysis, particularly when object boundaries are weak or ambiguous. Under such conditio
👤 Zihan Li, Jiebao Sun, Fanghui Song, Zhichang Guo 📅 2026-07-24 🔗 arXiv 📄 PDF
Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but underexplored factor affecting model generalisability is intensi
👤 Oliver Mills, Philip Conaghan, Samuel Relton 📅 2026-07-22 🔗 arXiv 📄 PDF
Current 3D generative models mostly produce a final surface: a visually strong but largely opaque mesh. Interactive 3D worlds need more than a surface. They need named parts, an assembly hierarchy, me
👤 Nimra Noor, Muhammad Bilal, Abdullah Hussain, Hassan Baig 📅 2026-07-22 🔗 arXiv 📄 PDF

🤖 Medical AI Agent & VLM

16条
Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is u
👤 Mateusz Kozłowski 📅 2026-07-28 🔗 arXiv 📄 PDF
Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require
👤 Zheng Tong, Yang Liu, Wanshu Fan, Jing Qin et al. (9 authors) 📅 2026-07-28 🔗 arXiv 📄 PDF
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must ab
👤 Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu et al. (24 authors) 📅 2026-07-27 🔗 arXiv 📄 PDF
Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large language model (LLM) agent can do the labor of reconstr
👤 Andreas Maier, Lucas Kachelriess, Siming Bayer, Yixing Huang et al. (7 authors) 📅 2026-07-24 🔗 arXiv 📄 PDF
Few-shot learning enables medical image segmentation models to adapt to new tasks using only a small number of labelled examples. However, adaptation performance depends strongly on which examples are
👤 Chenlan Zhao, Benny Wong, Timothy F. Lundberg, Ahmed M. Elsayed et al. (11 authors) 📅 2026-07-24 🔗 arXiv 📄 PDF
This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models have demonstrated significant potential in the fiel
👤 Bannapol Limanond, Masanori Suganuma, Takayuki Okatani 📅 2026-07-24 🔗 arXiv 📄 PDF
Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pretraining (CLIP) has advanced medical vision langua
👤 Haoyu Jiang, Ziping Cong 📅 2026-07-24 🔗 arXiv 📄 PDF
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know which internal units encode clinical findings or where that informa
👤 Farhad Nooralahzadeh, Lea Bogensperger, Christian Bluethgen, Michael Krauthammer 📅 2026-07-23 🔗 arXiv 📄 PDF
Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases. Here we show that Monte Carlo dropout, applied to
👤 Frederik Hauke, Patrick Wienholt, Christiane Kuhl, Dyke Ferber et al. (7 authors) 📅 2026-07-22 🔗 arXiv 📄 PDF
Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most existing pathology ben
👤 Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang et al. (9 authors) 📅 2026-07-21 🔗 arXiv 📄 PDF
Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completeness, and semantic faithfulness. We introduce Dobic
👤 Thanni Adewuyi, Angelica Obayi, Andem Aniekan, Samuel Okoko et al. (13 authors) 📅 2026-07-21 🔗 arXiv 📄 PDF
Attention and saliency heatmaps are widely used to explain medical Vision-Language Model (VLM) outputs on chest X-rays, yet whether they truly highlight the image evidence driving predictions has not
👤 Binesh Sadanandan, Vahid Behzadan 📅 2026-07-20 🔗 arXiv 📄 PDF
Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather than isolated factual queries. Such complex answers may be score
👤 Yuhao Liu, Cheng Zhao, Guanghui Yue 📅 2026-07-20 🔗 arXiv 📄 PDF
Test-time adaptation (TTA) aims to mitigate distribution shifts by adapting models with unlabeled target data at inference time. While TTA with vision-language models (VLMs) has shown promising result
👤 Lingrui Li, Nan Pu, Dong Zhao, Wenjing Li et al. (7 authors) 📅 2026-07-20 🔗 arXiv 📄 PDF
Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often compromised by frequent errors in visual interpr
👤 Jianqin Liu, Weiwei Cao, Wanxing Chang, Ruifeng Yuan et al. (10 authors) 📅 2026-07-24 🔗 arXiv 📄 PDF
Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- re
👤 Pranav Kaliaperumal, Manisha Kaliaperumal 📅 2026-07-22 🔗 arXiv 📄 PDF

🔄 Medical Image Registration

2条
In practical settings, medical image segmentation models are often developed with limited annotated data rather than fully labeled datasets. Training frequently begins in ultra-low labeled regimes whe
👤 Bahram Jafrasteh, Cheng Wan, Heejong Kim, Johannes C. Paetzold et al. (5 authors) 📅 2026-07-27 🔗 arXiv 📄 PDF
Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains chall
👤 Jiaming Feng, Xukun Zhang, Shahid Farid, Sharib Ali 📅 2026-07-20 🔗 arXiv 📄 PDF