Skip to content
Health information portal — no online consultationsCash on delivery · shipped within 3 to 7 business days
Riverside Dental Care

Dental Innovations

AI dental diagnostics: precision in early cavity detection

A 2025 umbrella review with meta-analysis published in PLOS ONE reported pooled diagnostic performance of 0.85 sensitivity and 0.90 specificity for artificial intelligence algorithms detecting dental caries across multiple imaging modalities.

AI dental diagnostics: precision in early cavity detection

The reported 95% confidence intervals were 0.83–0.93 for sensitivity and 0.85–0.95 for specificity. A separate 2025 PubMed meta-analysis covering 45 studies found individual AI model accuracy ranging from 41.5% to 98.6%.

That spread is not a minor statistical detail. It is the central fact behind any discussion of AI dental diagnostics accuracy for early cavity detection. Pooled averages describe the combined evidence, but they do not erase the differences between architectures, image types, lesion definitions, annotation methods, or validation protocols. An algorithm that performs well on one curated radiographic dataset may behave differently when it encounters images from another scanner, another clinic, or patients with different patterns of disease.

The figures describe classification of visual patterns on bitewing radiographs, periapical films, and intraoral camera images. In most systems, the algorithm produces a probability score, lesion flag, or highlighted region for a tooth surface within the captured field. It does not directly measure mineral density and does not independently establish whether restorative treatment is necessary. That distinction matters both for clinical interpretation and for evaluating automated caries detection software.

Early stage tooth decay identification is especially demanding because initial lesions may occupy a small area, have weak contrast, or overlap with normal anatomical variation. A system may identify a suspicious pattern without being able to determine the full clinical significance of that pattern. The value of the output is therefore tied to how it is integrated with examination, patient history, radiographic quality, and the dentist’s judgment.

The Mechanics of Convolutional Neural Networks in Dental Imaging

Convolutional Neural Networks, or CNNs, operate as layered feature extractors. An input image passes through successive convolutional filters that detect edges, gradients, contrast changes, and textural patterns. Early layers generally capture low-level visual features, such as pixel-intensity boundaries and local contrast differentials. Deeper layers combine those signals into more complex representations.

In dental imaging, those higher-order representations may correspond to disruptions at the enamel-dentin interface, radiolucency gradients consistent with demineralization, or morphological discontinuities on proximal surfaces. The network does not interpret an image in the same conceptual way as a dentist. It calculates patterns associated with labeled examples and uses those learned relationships to assign probabilities to new images.

For bitewing radiography, a CNN receives a two-dimensional grayscale matrix. The exact image dimensions depend on the acquisition system and preprocessing pipeline; the draft studies describe matrices commonly ranging from 1,024 × 1,024 to 2,048 × 2,048 pixels. With 8-bit images, each pixel can carry a grayscale value from 0 to 255. The network applies learned kernels across local regions of the matrix and generates feature maps that represent detected structures.

Pooling or other downsampling operations reduce spatial dimensions while retaining features considered useful for classification. Later layers combine the remaining information, and the final output may assign a probability to a tooth, a tooth surface, or a suspected lesion. Some systems classify the entire region; others are designed to localize the suspected area more precisely.

The distinction between classification and localization is clinically important:

  • Classification asks whether an image or tooth surface belongs to a category, such as carious or sound.
  • Detection identifies where a suspected abnormality is located within the image.
  • Segmentation assigns pixels to a lesion or anatomical region, producing a more detailed map.
  • Probability scoring expresses model confidence but does not automatically establish a definitive diagnosis.
Pooled sensitivity and specificity from the 2025 umbrella review were 0.85 and 0.90. These are aggregate values from heterogeneous datasets, not guarantees for any individual image or patient.

For intraoral camera images, the technical problem changes. Color images contain red, green, and blue channels rather than a single grayscale channel. The model must account for illumination, camera angle, saliva, tooth morphology, surface staining, plaque, restorations, and the visual differences created by image capture. Training data must include enough variation for the model to distinguish a potentially carious surface from an artifact or a benign change in appearance.

The visual signals also differ from those in radiographs. A bitewing model may learn associations involving radiolucency and contrast through enamel and dentin. An intraoral image model may rely more heavily on chromatic variation, surface texture, opacity, and lesion borders. A model trained exclusively on bitewing radiographs cannot simply be transferred to intraoral photographs and expected to retain its performance. The modality, labeling process, and intended clinical question all need to match.

This is one reason why the phrase “AI dental diagnostics” covers several technically different products. Two systems may both claim to support caries detection while analyzing different image types, producing different outputs, and being evaluated against different reference standards. Their accuracy figures should not be treated as interchangeable.

Comparative Performance: AI Sensitivity vs. Human Diagnostic Precision

Direct comparisons with practicing clinicians provide a useful reference point, although they must remain tied to the conditions of the individual study. In Cantu et al. (2020), a deep-learning model trained on bitewing radiographs reached 0.80 accuracy on a test set. Practicing dentists reading the same bitewings reached 0.71 accuracy. The difference was more pronounced for sensitivity: the model scored 0.75, while the dentists scored 0.36.

Diagnostic metricDeep-learning modelPracticing dentists
Accuracy0.800.71
Sensitivity0.750.36

Sensitivity measures the proportion of actual caries cases that a system correctly identifies. On this test set, the model detected 75% of the lesions classified as carious and missed 25%. The dentists’ sensitivity of 0.36 means they identified 36% of the relevant cases and missed 64% under the study conditions.

The 0.39 sensitivity gap is therefore clinically meaningful within that dataset. It illustrates one of the potential benefits of AI-assisted dental exams: an algorithm can draw attention to subtle or easily overlooked findings and create a more consistent second pass over an image. In a busy practice, that may be useful for review, quality assurance, or triage.

Accuracy, however, is not the same as sensitivity. A system can produce a high overall accuracy while still missing a clinically important subset of lesions, particularly when sound surfaces substantially outnumber carious ones. Specificity must be considered alongside sensitivity because a tool that flags too many sound surfaces can create unnecessary follow-up, patient anxiety, or pressure toward overtreatment.

The comparison also does not mean that the model replaced the dentists’ full clinical process. The study compared readings of the same bitewings under defined conditions. It did not establish that an algorithm should make treatment decisions independently, nor did it show that AI performs better than every clinician in every setting. Clinical examination includes information that may not be visible in a single image: symptoms, previous restorations, caries risk, fluoride exposure, diet, lesion activity, and changes observed over time.

Image quality is another variable. A model may struggle with overlapping structures, motion blur, underexposure, overexposure, positioning errors, or incomplete visualization of the contact areas. A dentist may also be limited by poor imaging, but can sometimes recognize that the image is inadequate and request another view. An automated output should not be treated as a substitute for that quality judgment.

Lesion location changes the task as well. Occlusal lesions on molars present different detection challenges from proximal lesions on premolars. Early enamel changes may be harder to classify than larger lesions with clearer radiographic expression. The reference standard also matters: labels based on clinical criteria, expert consensus, or histological assessment do not necessarily describe the same disease threshold.

For these reasons, a single comparative study supports a bounded conclusion: on that test set, the deep-learning model showed higher accuracy and sensitivity than the participating dentists. It does not justify the universal claim that AI outperforms clinicians for all caries detection on all bitewing images.

Analyzing the Variance in AI Accuracy Across Clinical Studies

The 41.5%–98.6% accuracy range reported in the 2025 PubMed meta-analysis of 45 studies reflects substantial methodological heterogeneity. The studies did not all evaluate the same clinical problem with the same data. Some used external validation datasets, while others relied on internal cross-validation. Some focused on bitewing radiographs; others included panoramic imaging or near-infrared transillumination. Lesion definitions ranged from ICDAS-based criteria to histological reference standards.

Those choices affect the apparent performance of the model. A system evaluated on images collected under one protocol may benefit from consistent exposure, positioning, and scanner characteristics. A model tested on images from a different source population faces a harder generalization problem. The same algorithm can therefore produce very different accuracy figures without any single number being mathematically incorrect.

Internal validation and external validation answer different questions:

  • Internal validation tests performance using data connected to the development dataset, often through a train-test split or cross-validation.
  • External validation tests the model on images from a distinct source population, institution, or acquisition environment.
  • Clinical deployment adds further variation, including workflow integration, operator behavior, image quality, and the characteristics of the patients actually seen by the practice.

A model can perform strongly under internal validation and still lose accuracy when scanner settings, patient demographics, lesion prevalence, or annotation practices change. That is why external validation is a central issue in evaluating artificial intelligence in dental imaging.

The 2026 systematic review in MDPI Diagnostics evaluated 28 AI caries detection studies and observed high overall risk of bias across the available evidence. Only 5 of the 28 studies used external validation datasets. The remaining studies reported internal performance measures without demonstrating the same level of testing on independent data.

SourceStudies evaluatedStudies with external validation
MDPI Diagnostics systematic review (2026)285 of 28
PubMed meta-analysis (2025)45Not reported uniformly
Only 5 of 28 studies in the 2026 systematic review used external validation. Internal performance figures should not automatically be read as deployable accuracy.

The 5-of-28 result is not a reason to dismiss AI research. It is a reason to interpret impressive numbers carefully. A high internal score can demonstrate that a model learned patterns in the development data. It does not, by itself, show that the model will maintain that performance in a different clinic or across a broader patient population.

The 7-study quantitative meta-analysis from PubMed (2025) reported mean sensitivity of 76%, with a 95% confidence interval of 65%–85%, and mean specificity of 91%, with a 95% confidence interval of 86%–95%. These intervals describe uncertainty around the pooled estimates. They are not a guarantee that every future study will fall inside the same range, and they should not be explained by study count alone.

The fact that a meta-analysis includes fewer studies does not, by itself, produce narrower confidence intervals. Interval width depends on several factors, including the amount of information contributed by the studies, variability between estimates, sample sizes, and the statistical model used. In this evidence base, the relevant point is simply that the reported pooled estimates and their intervals should be presented as calculated—not assigned a causal explanation that the supplied results do not establish.

The specificity estimates across the major meta-analyses should likewise be stated precisely. The reported pooled values are approximately 0.90 and 0.91, so it is more accurate to describe pooled specificity as ranging from 0.90 to 0.91 than to say that it exceeds 0.90 in every analysis.

At the lower bound of the 95% confidence interval for specificity in the 7-study analysis, specificity was 86%. The corresponding false-positive rate is approximately 14%, because the false-positive rate is calculated as 1 minus specificity. It is therefore not correct to say that false positives remain below 14% at that endpoint. The interval indicates that performance could be consistent with a false-positive rate of about 14% at the lower-specificity boundary.

That conversion matters in practice. A specificity of 90% means that approximately 10% of sound surfaces would be classified as positive under the relevant conditions. A specificity of 95% corresponds to approximately 5% false positives. These are population-level performance measures, not predictions for a particular patient, and the number of flagged surfaces will also depend on how many sound and diseased surfaces are present in the evaluated population.

The Role of Architecture: MobileNet-v3 and U-Net in Intraoral Analysis

A clinical evaluation deployed a deep-learning model incorporating MobileNet-v3 and U-Net architectures on 4,361 teeth represented in intraoral camera images. The reported metrics were 93.40% overall accuracy, 81.31% sensitivity, and 95.65% specificity.

These three figures describe different aspects of performance. Overall accuracy summarizes correct classifications across the evaluated dataset. Sensitivity describes how many carious cases the model detected. Specificity describes how many sound tooth surfaces it correctly classified as non-carious. A strong specificity does not cancel out missed lesions, and a higher sensitivity does not automatically mean that the model is clinically preferable if it generates an impractical number of false positives.

MobileNet-v3 functions as the classification backbone. Its design emphasizes computational efficiency through mechanisms such as depthwise separable convolutions and a reduced parameter burden compared with larger conventional CNNs. That makes this type of architecture suitable for applications where computing resources, storage, or response time may be constrained.

In an intraoral workflow, efficiency can matter because the system may need to process images near the time of capture. A smaller backbone may make integration with portable equipment or routine practice software more feasible. It does not, however, make the model accurate by definition. Efficient computation and diagnostic validity are separate properties. The architecture still depends on the quality, diversity, and labeling of the training data.

U-Net serves a different role. Its encoder-decoder structure uses skip connections to preserve spatial information while building a representation of the image. The encoder extracts increasingly complex features; the decoder reconstructs a detailed output that can identify the location of a suspected lesion. This makes U-Net appropriate for segmentation, where the goal is to mark pixels or regions rather than assign only a single label to the entire tooth.

The combined architecture can therefore support a more informative output:

  • MobileNet-v3 helps classify the visual content efficiently.
  • U-Net helps localize the suspected carious area.
  • The resulting image can show both the classification result and the region that influenced it.
  • The clinician can compare the highlighted area with the original photograph and the rest of the examination.

That last step is essential. A segmentation mask is not a histological map. It is the model’s estimate of which pixels resemble the examples used during training. A highlighted border may help focus attention, but it cannot establish lesion activity, depth, cavitation, or the appropriate treatment by itself.

The 93.40% accuracy figure applies to the combined classification-segmentation output across the 4,361-tooth dataset. The 81.31% sensitivity indicates that the model identified a substantial majority of carious cases under the study conditions, but it still missed some. The 95.65% specificity indicates that most sound surfaces were correctly excluded from the caries classification, while a smaller proportion were still flagged.

The asymmetry between sensitivity and specificity may be useful in a review workflow, but it should not be described as proof that the system has solved the balance between missed disease and unnecessary intervention. The appropriate threshold depends on the clinical purpose. A tool used to prompt a second look may tolerate a different false-positive rate from a tool intended to trigger a formal diagnostic pathway.

Clinical Limitations and the Future of AI-Assisted Decision Support

The 2026 MDPI Diagnostics review identified a high risk of bias across the evaluated studies. The concerns included non-representative training datasets, single-center image acquisition, retrospective labeling, and limited external validation. These factors can make performance metrics look more stable or more impressive than they will be in routine deployment.

A dataset assembled at one center may reflect that center’s imaging equipment, patient mix, referral patterns, and diagnostic habits. Retrospective labels may also encode the assumptions of the clinicians or reference process used to create them. If the labels are inconsistent, the algorithm may learn a noisy approximation of disease rather than a reliable clinical target.

Commercial performance varies for the same reasons. Different products may use different architectures, image modalities, lesion definitions, thresholds, and validation procedures. No universal accuracy benchmark applies to every automated caries detection system. The published range of 41.5%–98.6% demonstrates how strongly the result can depend on the study and evaluation design.

A practice considering automated software should examine the evidence behind the specific product rather than relying on the strongest number in the general literature. Relevant questions include:

  • Was the software evaluated on the same imaging modality used by the practice?
  • Was the test set independent of the training data?
  • Were images obtained from more than one site or acquisition environment?
  • How were caries and sound surfaces defined?
  • Are sensitivity and specificity reported separately, rather than only overall accuracy?
  • Does the system show the suspected region, or does it provide only a binary label?
  • What happens when the image quality is inadequate or the case falls outside the training distribution?

These questions are more useful than asking whether AI is accurate in the abstract. The practical issue is whether a particular system produces interpretable, appropriately calibrated assistance for a particular workflow.

AI-assisted decision support is best understood as an additional layer of review. Its output may be a probability map, a highlighted lesion, or a flag that directs the clinician back to a particular tooth surface. The dentist remains responsible for integrating that output with examination findings, radiographic interpretation, risk assessment, prior records, and the patient’s preferences.

The reported sensitivity and specificity values make the limits visible. At sensitivity levels of roughly 75%–85% across major pooled analyses, false negatives remain possible. At specificity levels of roughly 90%–95%, false positives also remain possible. Neither measure reaches 100%, and neither describes the complete clinical reasoning process.

The technology may be particularly useful where consistency is the goal: reviewing a large image set, prompting a second look at subtle findings, supporting communication with a patient, or creating a quality-assurance layer around radiographic interpretation. Its role is less defensible when the output is treated as an autonomous treatment recommendation or as proof that a lesion is active, cavitated, or ready for restoration.

AI-assisted caries detection can deliver measurable gains over unaided human reading in specific datasets, including the Cantu et al. comparison of 0.75 sensitivity for AI versus 0.36 for dentists. But pooled specificity is better described as approximately 0.90–0.91 across the major meta-analyses, not as a value that exceeds 0.90 in every estimate.

The most defensible position is neither uncritical enthusiasm nor blanket rejection. AI dental diagnostics accuracy for early cavity detection is promising, but it is conditional. Architecture, imaging modality, reference standard, patient population, and validation design all shape the result. A clinic evaluating these systems should give greater weight to vendor-specific external validation and transparent reporting than to a single pooled average.

Used in that way, AI can strengthen the diagnostic process without pretending to replace it. The software can point, compare, and prioritize. The clinical examination still has to decide what the finding means.

FAQ

How accurate is AI at detecting dental cavities?
Accuracy varies widely depending on the study, ranging from 41.5% to 98.6%. Pooled data from meta-analyses typically show sensitivity around 75%–85% and specificity around 90%–91%.
Does AI perform better than dentists at identifying caries?
In some comparative studies, such as the 2020 Cantu et al. analysis, AI models have demonstrated higher sensitivity than dentists under specific test conditions. However, these results are limited to those specific datasets and do not imply that AI outperforms clinicians in all clinical settings.
What is the difference between classification, detection, and segmentation in AI dental software?
Classification determines if a tooth surface is carious or sound, detection identifies the location of a suspected abnormality, and segmentation maps specific pixels to a lesion or anatomical region.
Why is external validation important for AI dental tools?
External validation tests a model on images from a different population or environment than the one used for training. This is critical because a model that performs well during internal testing may lose accuracy when applied to different scanners, clinics, or patient demographics.
Can AI replace a dentist's clinical examination?
No, AI is intended as an additional layer of review. It cannot account for clinical factors like patient symptoms, caries risk, fluoride exposure, or lesion activity, which require a dentist's professional judgment.