Ultrassom 360°
Artificial Intelligence14 min read

AI for detection, classification, and reporting: reducing inter-operator variability in ultrasound

Recent studies show that AI algorithms for detection, classification, and measurement in ultrasound can match or exceed average radiologist accuracy and consistently reduce inter-observer variability.

Published on September 12, 2026Last updated on September 12, 2026

Ultrasound is, by nature, the imaging modality most dependent on the operator. Unlike CT or MRI, where an automated protocol acquires a standardized volume regardless of who runs the machine, ultrasound requires the person holding the transducer to choose the scanning plane, adjust gain and depth, recognize anatomy in real time, and decide which image to capture and measure. This dependence on manual skill and instantaneous visual judgment is the main documented source of inter-operator and inter-observer variability, affecting everything from ejection fraction estimation to thyroid and breast nodule classification.

It is precisely this characteristic that makes artificial intelligence particularly relevant for ultrasound, even more so than for other modalities. An algorithm does not tire at the end of a shift, its attention does not vary with accumulated experience, and it applies the same criteria identically to the first and the thousandth exam. In recent years, a growing number of studies published in journals such as Radiology, European Heart Journal – Cardiovascular Imaging, Ultrasound in Obstetrics & Gynecology, and Gland Surgery have quantified, with concrete numbers for accuracy, area under the curve (AUC), and inter-observer agreement (kappa or ICC), the extent to which AI delivers on this promise.

Detection and classification of thyroid nodules

A systematic review with meta-analysis published in Gland Surgery, pooling 69 studies on AI systems for detecting malignancy in thyroid nodules on ultrasound, found a combined sensitivity of 89% (95% CI: 87-91%), specificity of 84% (95% CI: 80-88%), and AUC of 0.93 (95% CI: 0.91-0.95). In direct paired analysis against experienced radiologists across nine studies, AI achieved an AUC of 0.922 versus 0.842 for radiologists (statistically significant difference, p<0.0001), with similar sensitivity (85.2% vs. 85.8%) but markedly higher specificity (83.6% vs. 71.4%) — meaning AI is less likely to flag benign nodules as suspicious, reducing unnecessary biopsies.

A study published in Radiology on an AI model based on ACR TI-RADS categories (MTI-RADS) reported an AUC of 0.91 (95% CI: 0.87-0.95), sensitivity of 83%, and specificity of 87% for malignancy diagnosis. Compared with junior radiologists, the model performed significantly better on AUC (0.78, p<0.001), sensitivity, and specificity; compared with experienced radiologists, performance was equivalent on AUC and sensitivity, but AI still showed higher specificity (87% vs. 80%, p=0.02). This illustrates a recurring pattern in the literature: AI tends to level up the performance of less experienced examiners, acting as a second eye that narrows the gap between beginners and specialists.

Breast nodules and support for BI-RADS categorization

In breast ultrasound, commercial decision-support systems already have robust accumulated data. S-Detect (Samsung), evaluated in a meta-analysis of 11 studies, showed sensitivity of 0.82 (95% CI: 0.74-0.88), specificity of 0.83 (95% CI: 0.78-0.88), and AUC of 0.90. The Koios DS system achieved a mean AUC of 0.88 when used alone; interestingly, when radiologists used Koios as support (US + DS), the reader's mean AUC was 0.87, versus 0.83 for conventional ultrasound without support — a modest but consistent improvement in the human examiner's accuracy.

The most striking gain, however, usually appears in inter-observer agreement rather than in isolated accuracy. With Koios DS, inter-observer agreement (Kendall's tau-b) rose from 0.54 to 0.68, and intra-observer variability fell from 13.6% to 10.8% discordant cases. With the BU-CAD system, reader AUC increased from 0.758 to 0.829 with AI support, and reading time dropped by about 40% (from 30.15 to 18.11 seconds per case). For BI-RADS 4 lesions specifically — the category with the greatest variability and the highest number of biopsies with benign results — real-time AI systems have been studied specifically to reduce the rate of unnecessary biopsies without compromising cancer sensitivity.

Automated measurement and reduction of quantitative variability

Perhaps the most mature example of AI reducing quantitative variability comes from echocardiography. A study published in European Heart Journal – Cardiovascular Imaging validated an AI method for automatically measuring ventricular volumes and ejection fraction, comparing it to manual measurements in internal and external databases. The mean difference for ejection fraction ranged from -5.5 to 0.3 percentage points, with successful analysis in 95-100% of exams. The most relevant finding, however, was methodological: all 95% confidence intervals for the variance ratio fell below 1.0, indicating that AI variability was lower than inter-observer variability between two human cardiologists — and the method was also non-inferior to same-operator test-retest repetition.

In obstetrics, a study published in BMC Pregnancy and Childbirth evaluated an automated fetal biometry software (CUPID) in the second trimester. The AI's intra-observer reproducibility was virtually perfect (ICC = 1.0), compared with ICC of 0.974-0.999 for a senior radiologist and 0.749-0.934 for a junior radiologist. The AI's mean absolute error ranged from 0.36 to 2.53 mm across biometric measurements, while the junior radiologist showed errors of 0.64 to 8.13 mm — a striking difference precisely in the less experienced examiner group, where inter-observer variability most compromises gestational age and fetal weight estimation. The AI also processed each measurement in 0.05-0.07 seconds, versus 4.79-13.44 seconds manually.

Similar results have been described for carotid intima-media thickness (IMT), a classic marker of subclinical cardiovascular risk whose absolute value changes little between healthy and diseased individuals — making it extremely sensitive to small variations in manual measurement technique. Semi-automated and fully automated edge-detection systems using active contours, studied since the 2000s and more recently refined with neural networks (including attention/transformer architectures), aim precisely to eliminate the subjectivity of manually placing calipers at the intima-lumen and media-adventitia interfaces, identified as the main source of inter-operator variability in this type of measurement.

Real-time image-quality guidance

Beyond interpreting already-acquired images, a new generation of tools operates during acquisition itself, flagging in real time whether the scanning plane, framing, or gain are adequate before the image is saved and measured. A study on intra- and inter-observer agreement of an objective transvaginal image-quality scoring system, designed specifically to train AI algorithms, illustrates why this step matters: if input image quality already varies between operators, any automated classification or measurement built on it inherits part of that variability. Full acquisition automation — where the system itself guides or performs the scan — is covered in depth in another article in this series; here it is worth noting that ensuring input image quality is now recognized as a prerequisite for AI's accuracy gains in detection and classification to hold up in clinical practice.

AI-assisted report generation

A rapidly expanding field combines structured findings from detection/classification AI with language models to automatically generate report text. Systems described recently in the literature — including approaches that integrate radiologist annotations with deep learning analysis to synthesize structured multimodal breast ultrasound reports, and tools that turn the examiner's spoken descriptions during the exam into structured reports using large language models (LLMs) — aim to reduce typing time, standardize terminology (for example, ensuring that the descriptors used actually match the assigned BI-RADS or TI-RADS category), and reduce transcription errors. These systems remain mostly semi-automated: the AI proposes a structured draft from the imaging findings, but final validation of the text and clinical conduct remains with the responsible physician.

Direct evidence on inter-observer agreement

It is worth gathering, side by side, the agreement figures cited throughout this article, because that metric answers directly the clinical question about variability. For thyroid nodules, inter-observer agreement with the Koios DS system rose from 0.622 to 0.876 (agreement coefficient scale), and the AI-SONIC system achieved a Kendall coefficient of 0.995, exceeding the typical agreement between human examiners, which ranges from 0.72 to 0.85. For breast nodules with Koios DS, agreement rose from 0.54 to 0.68 (Kendall's tau-b). For ejection fraction, all of AI's variance confidence intervals fell below human inter-observer variability. Taken together, these data — obtained across different populations and equipment, with different software vendors — converge on the same conclusion: AI tends to measurably and reproducibly reduce the outcome variation that today depends on who performs or interprets the exam.

Honest limitations: bias, generalizability, and regulation

These results should not be read as a substitute for medical judgment. The first limitation is training-data bias: most commercial algorithms have been developed and validated predominantly on populations and equipment from a small number of centers, often in East Asia (for thyroid) or North America and Europe (for breast and heart), and their performance can drop when applied to populations with a different disease prevalence, body habitus, or equipment profile than the training data — a generalizability problem widely discussed in the literature on regulation of AI imaging devices. The second is equipment and configuration dependence: a model trained on images from one transducer manufacturer may lose accuracy on machines from another brand or with different gain/frequency settings.

From a regulatory standpoint, approval status varies by country, by intended use, and by software version. In the United States, dozens of imaging algorithms with an AI/machine-learning component have received FDA authorization in recent years, most through the 510(k) pathway based on prior predicates — which, according to reviews published in Radiology and regulatory-policy journals, raises questions about the transparency of validation data and about how well post-approval software updates continue to be adequately overseen. In Brazil, ANVISA is responsible for regulating software as a medical device (SaMD); before adopting any AI tool in clinical practice, physicians should verify it holds current registration for the intended indication in the country where they practice, not merely in another market.

Finally, the emerging consensus in the literature — reinforced by position papers on AI in ultrasound that discuss accountability and governance — is that these tools should be used as decision support, not as a replacement for medical evaluation. Even studies with the best accuracy results recommend that cases where AI and examiner disagree, incidental findings outside the algorithm's scope, and clinical management decisions (biopsy, follow-up, referral) remain under the responsibility and signature of a physician. AI reduces variability and can increase consistency, but final responsibility for diagnosis and management remains medical.

Content intended for health education and updates, and does not replace individualized medical evaluation.

← Back to articles