AiPhreaks ← Back to News Feed

Introducing NV-Reason-CT Open 3D CT VLM for Radiologist Chain-of-Thought Reasoning

By Jakub Antkiewicz

2026-09-24T13:14:10Z

NVIDIA has introduced NV-Reason-CT, an open research vision language model (VLM) specifically designed to interpret volumetric 3D computed tomography (CT) scans. The model addresses a significant gap where most modern VLMs fail to effectively process the data-dense nature of 3D medical imaging. By generating chain-of-thought reasoning that emulates a radiologist's systematic review process, the model aims to provide the auditability and trust required for clinical settings. NV-Reason-CT has already established a new state-of-the-art benchmark on the CT-RATE challenge with a Macro-F1 score of 0.614, outperforming existing 3D and fused 2D/3D models.

A Purpose-Built Architecture for Volumetric Data

Unlike standard VLMs that process 3D scans as a sequence of disconnected 2D images, NV-Reason-CT utilizes a full 3D architecture to preserve critical spatial context. This allows it to understand the shape, extent, and density of structures as they exist in three dimensions. The model's foundation is a combination of a dedicated 3D vision encoder and a powerful language model, trained end-to-end on a large cohort of CT data with structured reports and expert reasoning traces.

  • Full 3D ViT Encoder: The model employs a 3D vision transformer (ViT), adapted from Primus 3D ViT, which processes CT volumes as a true 3D input, preserving anatomical continuity between slices.
  • Language Model: It integrates a Qwen3.5-4B large language model, fine-tuned to generate structured reports and articulate diagnostic reasoning step-by-step.
  • Spatial Awareness: A technique called 3D MRoPE is used to ensure the language model understands the 3D spatial relationship between different parts of the scan throughout its reasoning process.
  • Two-Stage Training: The model is first trained via supervised fine-tuning on over 550,000 examples of radiologist reasoning data and then refined using Group Relative Policy Optimization (GRPO) to improve the clinical correctness of its outputs.

Building Trust in Medical AI

By releasing NV-Reason-CT as an open research foundation rather than a closed clinical product, NVIDIA is positioning it as a foundational layer for the broader medical AI ecosystem. The model's ability to produce transparent, step-by-step reasoning was validated for its clinical plausibility by radiologists at the National Institutes of Health (NIH). This emphasis on explainability is critical, as it allows clinicians to understand and verify the AI's conclusions, a key step toward integrating such tools safely into diagnostic workflows. The open model supports post-training, enabling developers to create specialized applications for specific CT analysis tasks.

By releasing NV-Reason-CT as an open research foundation, NVIDIA is not just pursuing state-of-the-art performance but is strategically targeting the primary barrier to AI adoption in clinical settings: trust. The model's emphasis on emulating radiologist chain-of-thought reasoning is a direct attempt to make AI outputs auditable and explainable, transforming diagnostic tools from black-box classifiers into interactive partners and lowering the barrier for developing specialized, regulatory-compliant applications.
End of Transmission
Scan All Nodes Access Archive