Conceptio › Archive › NCBI PubMed Central
NCBI PubMed Centralopen access

Automatic estimation of single-tooth width from standardized two-dimensional occlusal photographs using deep learning.

Li L et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed systems architecture

Automatic estimation of single-tooth width from standardized two-dimensional occlusal photographs using deep learning - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Head Face Med . 2026 Mar 9;22:41. doi: 10.1186/s13005-026-00605-1 Search in PMC Search in PubMed View in NLM Catalog Add to search Automatic estimation of single-tooth width from standardized two-dimensional occlusal photographs using deep learning Litong Li Litong Li 1 Department of Orthodontics, School of Stomatology, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration , National Clinical Research Center for Oral Diseases ,Shaanxi Clinical Research Center for Oral Diseases, The Fourth Military Medical University, Xi’an, Shaanxi 710032 China Find articles by Litong Li 1, # , Yulin Shi Yulin Shi 2 Department of Maxillofacial Trauma and Orthognathic Surgery, School of Stomatology, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration , National Clinical Research Center for Oral Diseases ,Shaanxi Clinical Research Center for Oral Diseases, The Fourth Military Medical University, Xi’an, Shaanxi 710032 China Find articles by Yulin Shi 2, # , Duo Wang Duo Wang 3 School of Computer Science, Northwestern Polytechnical University, Xi’an, 710072 China Find articles by Duo Wang 3 , Weixu Li Weixu Li 1 Department of Orthodontics, School of Stomatology, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration , National Clinical Research Center for Oral Diseases ,Shaanxi Clinical Research Center for Oral Diseases, The Fourth Military Medical University, Xi’an, Shaanxi 710032 China Find articles by Weixu Li 1 , Tingting Liu Tingting Liu 1 Department of Orthodontics, School of Stomatology, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration , National Clinical Research Center for Oral Diseases ,Shaanxi Clinical Research Center for Oral Diseases, The Fourth Military Medical University, Xi’an, Shaanxi 710032 China Find articles by Tingting Liu 1 , Yuanqing Ma Yuanqing Ma 1 Department of Orthodontics, School of Stomatology, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration , National Clinical Research Center for Oral Diseases ,Shaanxi Clinical Research Center for Oral Diseases, The Fourth Military Medical University, Xi’an, Shaanxi 710032 China Find articles by Yuanqing Ma 1 , Meng Cao Meng Cao 1 Department of Orthodontics, School of Stomatology, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration , National Clinical Research Center for Oral Diseases ,Shaanxi Clinical Research Center for Oral Diseases, The Fourth Military Medical University, Xi’an, Shaanxi 710032 China Find articles by Meng Cao 1, ✉ Author information Article notes Copyright and License information 1 Department of Orthodontics, School of Stomatology, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration , National Clinical Research Center for Oral Diseases ,Shaanxi Clinical Research Center for Oral Diseases, The Fourth Military Medical University, Xi’an, Shaanxi 710032 China 2 Department of Maxillofacial Trauma and Orthognathic Surgery, School of Stomatology, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration , National Clinical Research Center for Oral Diseases ,Shaanxi Clinical Research Center for Oral Diseases, The Fourth Military Medical University, Xi’an, Shaanxi 710032 China 3 School of Computer Science, Northwestern Polytechnical University, Xi’an, 710072 China ✉ Corresponding author. # Contributed equally. Received 2025 Aug 26; Accepted 2026 Mar 2; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13081552  PMID: 41796356 Abstract Objective To develop a deep learning-based fully automated model for estimating the mesiodistal width of teeth in standardized occlusal photographs, enabling accurate and efficient assessment of individual tooth dimensions. Methods The dataset comprised 14,403 teeth, including 12,079 teeth from an internal dataset and 2,328 teeth from an independent external test set. Clinically obtained mesiodistal crown widths derived from three-dimensional (3D) intraoral scans served as reference values. Corresponding mesial and distal keypoints were annotated on two-dimensional (2D) intraoral photographs. A two-stage deep learning framework was developed, consisting of tooth keypoint detection followed by a depth-informed regression model to estimate mesiodistal crown widths. Model performance was evaluated using key metrics, including mean absolute error (MAE), root mean squared error (RMSE), and the coefficient of determination (R²). Results The deep learning model demonstrated high accuracy across 12,079 teeth: MAE = 0.34 mm (95% confidence interval [CI]: 0.33–0.34); RMSE = 0.44 mm; R² = 0.94. External validation on 2,328 independent teeth demonstrated consistent performance (MAE = 0.33 mm, RMSE = 0.42 mm). Conclusions Our deep learning approach reliably estimates mesiodistal crown widths from standardized occlusal photographs, demonstrating strong agreement with 3D intraoral scan reference measurements and good clinical reproducibility across independent datasets within the same standardized workflow. Future studies will further investigate its robustness across heterogeneous imaging devices and acquisition conditions. Supplementary Information The online version contains supplementary material available at 10.1186/s13005-026-00605-1. Keywords: Deep learning, Intra-oral photograph, Artificial intelligence, Tooth width, Dental morphology Introduction Dental morphology serves as a fundamental cornerstone of modern stomatology. Of the many parameters involved in dental morphology, the mesiodistal width of an individual tooth constitutes a critical anatomical measurement with profound clinical significance. This metric is indispensable for orthodontic diagnosis and treatment planning, such as space analysis, the assessment of crowding, and Bolton ratio analysis [ 1 – 3 ], providing essential guidance for the design of prosthodontic dimensions and the evaluation of implant space [ 4 ], and serves as a baseline indicator for assessing dental developmental anomalies, the severity of attrition, and epidemiological studies [ 5 ]. Current measurement methodologies primarily involve manual caliper measurements on plaster models or three-dimensional (3D) software analyses of intraoral scans [ 6 ]. However, these approaches are associated with notable limitations: operational complexity and time consumption, inter- and intra-observer variability, cumbersome storage and transportation of plaster models [ 7 ], and the restricted accessibility of 3D scanning systems due to high procurement/maintenance costs and technical sensitivity [ 8 ]. Furthermore, these methods introduce additional steps into clinical workflows, thus compromising efficiency. Furthermore, 3D scan data are susceptible to loss or corruption over time due to software updates, hardware obsolescence, or file management issues, which may render valuable clinical records irretrievable. Consequently, these methods introduce additional steps, risks of data unavailability, and potential inefficiencies into clinical workflows. Over recent years, two-dimensional (2D) intraoral photographs have emerged as one of the most prevalent documentation tools in dental practice and research, owing to their advantages of rapid acquisition, cost-effectiveness, operational simplicity, and ease of storage/sharing [ 9 ]. Consequently, leveraging this widely available 2D photographic resource to achieve automated and high-precision dental measurements holds substantial practical significance and clinical value. This advancement could significantly enhance workflow efficiency, minimize human error, and facilitate large-scale data analysis. The automated analysis of 2D medical images using computer vision and artificial intelligence (AI) technologies [ 10 – 15 ] has emerged as a major research hotspot over recent years. In the field of dentistry, previous studies have investigated deep learning-based approaches for tasks such as the automated classification of intraoral photographs [ 16 ], the detection of caries [ 17 – 19 ], and the recognition of abnormal dentition [ 20 ]. Concurrently, several commercial AI-assisted dental monitoring platforms have been developed, primarily focusing on treatment tracking and progress assessment. With regards to dental measurements, some investigations have attempted to localize dental landmarks or segment tooth contours in photographs [ 21 , 22 ], thus providing potential foundations for subsequent dimensional analysis. However, distance prediction methods based on 2D imaging typically require specialized hardware (e.g., depth cameras or laser scanners) and precise camera parameters, which are often unavailable during the routine acquisition of intraoral photographs. Consequently, achieving high-precision and fully automated estimatio n of single-tooth width specifically from 2D intraoral photographs as an independent and primary research objective remains challenging and has yet to be investigated in detail. Although previous studies have attempted to derive more complex orthodontic model analysis indicators from intraoral photographs [ 21 , 22 ], most rely on simplified assumptions, such as estimating full dentition dimensions through proportional relationships of specific teeth, which may compromise the accuracy of individualized measurements. However, clinical decision-making requires precise anatomical data at the single-tooth level, particularly when dealing with non-standard dentitions, such as prostheses and developmental anomalies. Therefore, the development of direct measurement methods that are independent of prior proportional relationships has become a critical prerequisite for enhancing the reliability of digital oral analysis. Strengthening the measurement accuracy of fundamental parameters serves as the foundation for achieving higher-level automated analyses. In this study, we address these challenges by proposing a fully automated deep learning framework for estimating single-tooth mesiodistal width from 2D intraoral photographs. The framework employs a Convolutional Block Attention Module-High-Resolution Network (CBAM-HRNet) for keypoint localization and an innovative Tooth Distance Estimation Network (TDENet) to regress width from image features, including depth-informed structural cues. Rather than aiming to replace established 3D modalities, this work seeks to unlock the quantitative potential of the vast existing and routinely acquired archives of 2D intraoral photographs. By providing automated, clinically acceptable estimations from these ubiquitous yet underutilized records, the method enables retrospective epidemiological studies, facilitates screening and initial assessment in resource-conscious settings, and offers a practical tool for cases where 3D scans are unavailable or impractical. Our approach is validated against 3D scan-derived references and is designed for use with standardized occlusal photographs, offering a cost-effective and efficient complement to conventional measurement techniques. Materials and methods This study was conducted retrospectively and was approved by the Medical Ethics Committee of the Third Affiliated Hospital of Air Force Medical University on 3 April 2025(Approval No. KQ-YJ-2025-127). Written consent was waived as the study used anonymized retrospective data. Data collection We first obtained raw orthodontic records from two clinical centers within our institution. Between June 2024 and March 2025, the Department of Orthodontics at the Third Affiliated Hospital of Air Force Medical University provided 3,594 consecutive cases. From March 2025 to June 2025, the hospital’s Secondary Treatment Center supplied an additional 178 cases. All intraoral 3D scans were acquired using a TRIOS 3 (3shape, Copenhagen, Denmark) intraoral scanner. Intraoral photographs were captured with a digital camera (32.5-megapixel, Canon EOS 90D, Canon Inc., Japan) equipped with a macro single-lens reflex (SLR) lens (EF 100 mm f/2.8 L Macro IS USM, Canon Inc., Japan). The imaging setup included plastic retractors and intraoral mirrors to ensure optimal visibility. The camera settings were as follows: manual mode (M), auto white balance, ISO 200, shutter speed 1/60 s, aperture F32.0, auto flash exposure compensation, auto light exposure compensation, and ring flash activation. Data screening The inclusion criteria were as follows: aged 10–50 years; availability of both intraoral scans and corresponding clinical photographs at initial diagnosis; complete intraoral scan data (including full maxillary and mandibular dentition); and clear clinical photographs (appropriate exposure, accurate focus, without blurring, saliva interference, or soft tissue obstruction of key dental positions). The exclusion criteria were as follows: non-initial visit cases (e.g., follow-up or postoperative evaluations); technical defects in intraoral scans or photographs (e.g., scan discontinuities, blurred or overexposed images); and patients with tooth defects due to caries, trauma, or other causes. All orthodontic records from both the Main Department and the Secondary Treatment Center were rigorously screened against the predefined inclusion and exclusion criteria detailed above. Eligible patients were then assigned to either the internal dataset or the independent external test set based on their source clinic. Additionally, to further assess model robustness to changes in tooth position, we selected a longitudinal cohort from the Secondary Treatment Center. This cohort consisted of 15 patients who completed treatment without interproximal enamel reduction and for whom standardized occlusal photographs were available at both the pre-treatment (T0) and post-treatment (T1) stages. Data annotation First, we established gold standard measurements. Mesiodistal crown diameters were manually measured by an experienced orthodontist using Mimics Research 21.0 (Materialise, Leuven, Belgium). To assess intra-observer reliability, a random function was employed to select 1560 teeth from 65 patients for remeasurement after a 2-week interval. Analysis achieved a high level of intra-observer reliability (ICC = 0.999). Next, we performed intraoral photograph annotation. In accordance with the latest Guidelines for the Registration and Review of Artificial Intelligence Medical Devices [ 23 ], a three-tier annotation system was implemented in this study. The primary annotators were two master’s degree candidates with two years of clinical orthodontic experience and were responsible for the preliminary localization and annotation of the mesial and distal points of each tooth in intraoral photographs using LabelMe software (MIT, USA). The secondary reviewers consisted of two attending orthodontists with 5–10 years of clinical experience, tasked with verifying, correcting, and confirming data with positional discrepancies identified during primary annotation. The tertiary arbitrator was one chief orthodontist with over 10 years of clinical expertise, responsible for adjudicating cases unresolved by the reviewers (Fig. 1 ). Fig. 1. Open in a new tab Example Annotation. A The mesiodistal diameter of tooth crowns was manually measured in three-dimensional files using Mimics Research 21.0 software. B The mesial and distal points of teeth were annotated using LabelMe software Model development and establishment We propose a two-stage framework to regress 3D tooth widths from 2D photographs. The approach is motivated by projective geometry: the depth difference between mesial and distal landmarks invalidates traditional single‑scale calibration (see Supplementary Material, Section  2). Stage one detects keypoints using CBAM-HRNet. In stage two, the Tooth Distance Estimation Network (TDENet) uses these keypoints as anchors to extract depth‑aware features, which are fed to an MLP for calibration‑free width prediction. Tooth keypoint detection model To detect the mesiodistal keypoints of teeth in intraoral images, we developed a high-resolution network (HR-Net) framework incorporating a spatial-channel attention mechanism. HR-Net has demonstrated superior performance for dense prediction tasks, including pose estimation and facial keypoint localization [ 24 – 27 ], owing to its parallel multi-resolution subnetworks and persistent high-resolution feature representation. Given the alignment between these task characteristics and the detection of keypoints in teeth, HR-Net was selected as the backbone network. This network achieves keypoint localization via heatmap regression. Considering that tooth keypoints exhibit specific spatial structures and relative distance patterns in intraoral images, a convolutional block attention module (CBAM) was integrated into the network. By leveraging its dual attention mechanism (channel-wise and spatial), CBAM enhances feature representation, thereby improving the ability of the model to accurately capture spatial relationships between maxillary and mandibular teeth. The modified model (CBAM-HRNet), as illustrated in Fig. 2 , significantly improved the localization accuracy of tooth keypoints while maintaining high-resolution feature representation. Fig. 2. Open in a new tab Architecture of the CBAM-HRNet model During training, random horizontal flipping ( p = 0.5 with synchronous key-point mapping), color jittering (random perturbations of brightness, contrast, saturation, and hue), and random cropping (scaling to 80%–100% of the original area) are employed to enhance model symmetry and improve robustness to illumination and viewpoint variations. The Adam optimizer was adopted for training with an initial learning rate of 0.0002 and a weight decay of 0.01. A cosine annealing learning rate scheduler was utilized, and the model was trained for 200 epochs. Training was conducted on a GPU server with 24 GB of VRAM, with the batch size set to 24. Distance estimation model for dental keypoint pairs Following the localization of dental keypoints, the mesiodistal crown width of each tooth was estimated using a deep learning–based regression framework rather than being directly measured through geometric calibration. Specifically, the pixel-wise displacement between the predicted mesial and distal keypoints was treated as an intermediate spatial descriptor, which was subsequently non-linearly mapped to clinically measured mesiodistal widths through supervised learning. Considering that depth-related visual cues can enhance the discrimination of spatial relationships in intraoral images, we developed a Tooth Distance Estimation Network (TDENet) that integrates depth-informed structural features with 2D spatial representations. Importantly, no explicit camera calibration or physical scale reference was applied in this study. The depth information was incorporated solely to provide relative structural context and was not intended to recover true anatomical 3D geometry. The depth branch of TDENet utilized ZoeDepth, a monocular depth estimation model pre-trained on large-scale datasets, to extract relative depth features from intraoral photographs via transfer learning. In parallel, a spatial branch based on an encoder–decoder architecture captured 2D spatial distributions of dental keypoints. Multi-level features from both branches were progressively fused to form depth-aware joint representations, which encode spatial relationships in a learned feature space rather than explicit metric coordinates. These joint representations were then fed into a multilayer perceptron (MLP) regression head to predict mesiodistal crown widths. The overall framework of TDENet is illustrated in Fig. 3 . Fig. 3. Open in a new tab Architecture of TDENet. Note: The depth estimation model refers to the ZoeDepth model Given the sensitivity of depth-informed features to keypoint localization accuracy, minor deviations in annotated keypoint positions may lead to unstable feature representations. To improve robustness, an area-weighted probability-based keypoint perturbation augmentation strategy was introduced (Fig. 4 ). Within the tooth crown region, random perturbations were applied to keypoint P 1 , satisfying the constraints ||x-P 1 ||≤r and ||x-P 2 ||≥||x-P 1 ||, where P 2 denotes the nearest neighboring keypoint. This constraint ensured that perturbations remained within the anatomical neighborhood of P 1 while preserving relative spatial relationships, thereby guiding the model to learn smooth and stable depth-informed features and mitigating performance degradation caused by annotation variability. Collectively, this framework is designed to regress tooth widths from depth‑informed contextual features rather than from raw pixel distances. Fig. 4. Open in a new tab Schematic diagram depicting random keypoint offset The TDENet was built upon a transfer learning framework, utilizing the multi-scale depth features generated by the ZoeDepth model as input. During training, dental distance values served as supervision labels, and the parameters of relative depth structure provided by ZoeDepth were frozen; only the metric depth module was optimized. Differential hyperparameters were employed for distinct model components: the learning rates for the 2D network and metric depth module were set to 0.0002, while the MLP-Head component was set to 0.002. The AdamW optimizer was adopted for optimization, coupled with a cosine annealing learning rate scheduling strategy. Training was conducted on a GPU server (NVIDIA GeForce RTX 4090 with 24 GB of VRAM) for 200 epochs, with a batch size of 20. To precisely fit the lengths of 12 teeth, the Conflict-Averse Gradient (CAGrad) [ 28 ] algorithm, a gradient-based multi-task optimization method, was introduced. In addition, dynamic weight averaging was employed to adaptively adjust task loss weights, thus optimizing training efficacy and enhancing model generalizability. Further mathematical details regarding the keypoint perturbation strategy can be found in Supplementary Material, Section 1. Evaluation metrics To comprehensively evaluate model performance, we employed a proprietary dataset of 1,012 intraoral images for 5-fold cross-validation, with metrics averaged across test folds. An independent external test set of 194 images from a secondary treatment center further assessed reproducibility. The performance of the dental keypoint detection model was evaluated using the following metrics: normalized mean error (NME), percentage of correct keypoints (PCK), and area under curve (AUC, calculated based on the PCK curve). To evaluate the performance of the regression model for dental keypoint distance prediction tasks, three metrics were employed: mean absolute error (MAE), root mean squared error (RMSE), and coefficient of determination (R²). Specifically, the MAE was used to quantify the average absolute deviation between predicted and ground truth values, offering excellent interpretability. The RMSE demonstrates higher sensitivity to larger errors, thereby facilitating the identification of significant outliers in predictions. R² explains the explanatory power of the model with regard to the variance of ground truth values, thus serving as a crucial indicator of fitting performance. Statistical analysis Statistical analysis was performed using SPSS Statistics 26.0 software (IBM, USA). The Shapiro-Wilk normality test was applied to assess the distribution of dental data in each group. If the data followed a normal distribution, a paired t -test was employed to analyze the differences between the true values and model-predicted values. For non-normally distributed data, the Wilcoxon signed-rank test was used, with the significance level set at α = 0.05. Multiple testing correction was applied for the independent tests of 24 teeth. Finally, a two-way mixed-effects model was utilized to calculate the intraclass correlation coefficient (ICC [ 1 , 3 ]) to evaluate the clinical agreement between the true values and model-predicted values across different tooth groups, and the correlation coefficient was computed. Results Dataset and model development Following the screening process, the final study cohort comprised 506 eligible patients (193 males, 313 females; mean age 20.0 ± 7.9 years) from the Main Department, forming the internal dataset, and 97 qualified patients (38 males, 59 females; mean age 20.3 ± 8.0 years) from the Secondary Treatment Center, forming the independent external test set. The combined dataset included a total of 1,206 standardized occlusal photographs (603 maxillary and 603 mandibular views) with paired 3D intraoral scan files, all of which met the study’s technical quality requirements. The deep learning model for dental landmark detection was developed exclusively on the internal dataset. A five-fold cross-validation strategy was employed at the patient level to ensure robustness and mitigate overfitting. This approach partitioned the 506 patients into five folds, iteratively using four folds for training and one for validation. The final internal performance metrics represent the aggregate mean outcomes across all validation folds. The external test set remained completely blinded throughout this entire development and tuning process, serving solely for the ultimate evaluation of model stability. Evaluation metrics for the tooth landmark detection model The landmark detection pipeline was both highly efficient and accurate. The entire process from image input to output of all 12 tooth widths was completed in 0.075 s per arch. To evaluate detection accuracy, The mean width of a single central incisor (i.e., the average width of two central incisors) was adopted as the normalization factor. Model performance was comprehensively assessed using NME, PCK and AUC metrics, with detailed quantitative results presented in Table 1 . Table 1. Performance evaluation of the tooth landmark detection model NME [email protected] [email protected] [email protected] [email protected] Internal CV( n = 506) Maxilla 0.0599 0.940 1.000 0.250 0.400 Mandible 0.0795 0.581 0.887 0.107 0.240 External Test( n = 97) Maxilla 0.0603 0.928 0.948 0.295 0.425 Mandible 0.0801 0.598 0.938 0.082 0.221 Open in a new tab Internal results are presented as the mean value across the five folds; CV Cross-Validation; [email protected] represents the proportion of predicted landmarks within 8% of the central incisor width from ground truth positions; [email protected] represents the proportion of predicted landmarks within 10% of the central incisor width from ground truth positions; [email protected] represents the area under the PCK curve for error thresholds between 0–8%; [email protected] represents the area under the PCK curve for error thresholds between 0–10% In maxillary images, the internal dataset showed a normalized mean error (NME) of 0.0599, corresponding to an average pixel error of 3.6 pixels in 512 × 512 resolution images (central incisor width ≈ 60 pixels). This high accuracy was strongly corroborated on the external test set, which exhibited a nearly identical NME of 0.0603. For mandibular images, the internal set had an NME of 0.0795 (approx. 3.2 pixels, central incisor width ≈ 40 pixels), while the external set showed a virtually equivalent NME of 0.0801. The percentage of correct keypoints (PCK) metrics under strictly normalized scales revealed consistently strong performance across datasets. For the maxilla, the internal [email protected] reached 0.940 and [email protected] achieved 1.000. The external test set confirmed this robustness, with a [email protected] of 0.928 and [email protected] of 0.948. For the mandible, the internal [email protected] was 0.581 and [email protected] was 0.887. A very positive trend was observed externally, with PCK values slightly improving to 0.598 and 0.938, respectively. The performance advantage of incorporating the CBAM module is substantiated by a direct comparison with the baseline HRNet. Quantitative results of this ablation study are provided in Supplementary Tables S1, where CBAM-HRNet shows measurable gains in normalized mean error and percentage of correct keypoints. Evaluation of the distance estimation model for dental key points Overall model performance The distance-estimation model demonstrated high overall accuracy. On the internal dataset ( n = 12,079 teeth), it achieved an MAE of 0.34 mm and an ICC [ 1 , 3 ] of 0.98, confirming excellent agreement with the gold standard (Table 2 ). Importantly, this precision generalized to the independent external test set ( n = 2,328 teeth), where the MAE was 0.33 mm and the ICC was 0.99 (Table 3 ). The ME on the external set was − 0.05 mm, indicating minimal and clinically insignificant systematic bias. These results support the model’s consistent performance and clinical applicability within a standardized photographic workflow. Table 2. Performance of the Distance Estimation Model by Tooth Group on the Internal Dataset Tooth Group ME MAE RMSE R ² r ICC Maxilla Incisor -0.019 0.35 0.46 0.79 0.89 0.94(0.93–0.94) Canine -0.018 0.34 0.44 0.62 0.79 0.87(0.85–0.88) Premolar -0.006 0.30 0.40 0.71 0.84 0.90(0.90–0.92) Molar -0.004 0.43 0.54 0.38 0.62 0.72(0.68–0.75) Mandible Incisor -0.023 0.27 0.34 0.73 0.86 0.92(0.91–0.93) Canine -0.017 0.35 0.45 0.38 0.62 0.73(0.70–0.76) Premolar -0.036 0.32 0.41 0.69 0.83 0.90(0.89–0.91) Molar -0.014 0.43 0.56 0.5 0.71 0.81(0.78–0.83) Overall a -0.020 0.34 0.44 0.93 0.97 0.98(0.98–0.98) Open in a new tab ME Mean error, MAE Mean absolute error, RMSE Root mean square error, R ² coefficient of determination, r Pearson correlation coefficient, ICC Intraclass correlation coefficient a A total of 12,079 teeth were included Table 3. Performance of the distance estimation model by tooth group on the external test set Tooth Group ME MAE RMSE R ² r ICC Maxilla Incisor -0.049 0.35 0.46 0.80 0.89 0.94(0.93–0.95) Canine -0.055 0.35 0.45 0.65 0.81 0.88(0.85–0.91) Premolar -0.073 0.27 0.37 0.92 0.96 0.98(0.97–0.98) Molar 0.084 0.40 0.48 0.29 0.54 0.69(0.59–0.77) Mandible Incisor -0.072 0.26 0.33 0.59 0.77 0.85(0.82–0.88) Canine -0.068 0.33 0.39 0.30 0.55 0.66(0.55–0.74) Premolar -0.096 0.31 0.39 0.88 0.94 0.96(0.95–0.97) Molar -0.081 0.44 0.51 0.39 0.63 0.75(0.67–0.82) Overall a -0.045 0.33 0.42 0.94 0.97 0.99(0.98–0.99) Open in a new tab a A total of 2,328 teeth were included To assess agreement at the most granular level, we performed paired t-tests for each of the 24 tooth types, applying a strict Bonferroni correction ( α = 0.05/24). No differences reached significance on the external test set. In the much larger internal set, one tooth type differed significantly at this threshold, but the effect size was negligible (Cohen’s d = 0.16), indicating a statistically significant difference driven by sample size alone, without clinical relevance. Detailed results are provided in Supplementary Tables S2 and S3. Bland–Altman plots and regression analyses for the external test set are shown in Figs. 5 and 6 ; those for the internal dataset, which exhibited nearly identical agreement, are provided in Supplementary Figures S1 and S2. Fig. 5. Open in a new tab Analysis of maxillary tooth width estimation agreement on the external test set. A Regression plot of model-predicted versus gold-standard measurements. B Bland-Altman plot assessing the agreement between the two methods. The Y-axis shows the difference of predicted minus gold-standard values. The red line indicates the mean bias, and the dashed lines represent the 95% limits of agreement. The vast majority of data points fall within these limits, demonstrating excellent clinical agreement. ( N = 1,164). Fig. 6. Open in a new tab Analysis of mandibular tooth width estimation agreement on the external test set. A Regression and ( B ) Bland-Altman plots (see Fig. 4 for detailed method description). The model exhibited similarly high precision for mandibular teeth, with minimal bias and tight limits of agreement, confirming its robust performance across both arches. ( N = 1,164) Accuracy variations by tooth position Analysis revealed generally consistent performance patterns across tooth positions, though the external set showed more pronounced variations (Tables 2 and 3 ). Premolars and incisors maintained high accuracy in both cohorts, with maxillary premolars demonstrating particularly excellent ICCs of 0.90 (internal) and 0.98 (external). In contrast, molars and mandibular canines showed greater variability on the external test set. While these groups showed moderate internal ICCs (0.72–0.81), their external ICCs were substantially lower (maxillary molar = 0.69; mandibular canine = 0.66; mandibular molar = 0.75). This pattern is attributable to the greater susceptibility of posterior and lateral teeth to suboptimal camera angles and lens distortion (Fig. 7 ). Fig. 7. Open in a new tab Example of a dental photograph with inadequate canine exposure leading to landmark detection challenges. The image shows a maxillary canine with only the mesial incisal ridge visible; the distal ridge is obscured by a suboptimal camera angle. This common acquisition issue directly increases estimation error for canine teeth Robustness evaluation: model consistency across tooth position changes To preliminarily investigate the model’s prediction stability under changes in viewpoint, we selected 15 patients(4 males, 11 females; mean age 18.7 ± 8.0 years) from the second treatment center who had not undergone interproximal enamel reduction and had completed treatment. Standardized occlusal photographs at T0 (pre-treatment) and T1 (post-treatment) were obtained (15 maxillary and 14 mandibular occlusal images). All test images in this cohort were entirely novel to the model, having not been used for training, validation, or parameter tuning. For the same tooth, its true anatomical width remained constant between T0 and T1; however, due to orthodontic tooth movement, its projection in the two-dimensional images underwent significant changes. The trained model described above was used to predict dental element widths in both T0 and T1 photographs, and the consistency between these predictions was evaluated. The analysis results (Table 4 ) indicated a high level of consistency between the T0 and T1 predicted values across different tooth groups. The mean difference between T0 and T1 predictions was small (MD = 0.04 mm, SD = 0.26 mm). The Pearson correlation coefficients and ICC for most tooth groups were at a high level (overall ICC = 0.99, 95% CI: 0.99–0.99), suggesting that the predicted values maintained good correlation and absolute agreement even when dental positions changed. Molars and mandibular canines again showed a modest, though expected, decrease in performance on the external test set. This observation is consistent with their known sensitivity to imaging parameters. Nevertheless, their ICCs held at 0.73–0.89, remaining within the acceptable range. Table 4. Consistency of model predictions before and after orthodontic treatment Tooth Group MD ± SD r ICC Maxilla Incisor 0.08 ± 0.38 0.90 0.95(0.91–0.97) Canine 0.01 ± 0.14 0.60 0.73(0.43–0.87) Premolar 0.01 ± 0.22 0.99 0.99(0.99–0.99) Molar 0.03 ± 0.18 0.78 0.88(0.74–0.94) Mandible Incisor 0.04 ± 0.25 0.97 0.99(0.98–0.99) Canine 0.04 ± 0.20 0.81 0.89(0.77–0.95) Premolar 0.07 ± 0.26 0.99 0.99(0.99–0.99) Molar 0.04 ± 0.29 0.72 0.84(0.65–0.92) Overall a 0.04 ± 0.26 0.99 0.99(0.99–0.99) Open in a new tab MD Mean difference (T0 prediction-T1 prediction), SD Standard deviation of the paired differences a A total of 318 teeth were included Further comparison of the error distribution between model predictions and ground truth values revealed that it remained relatively stable between T0 and T1. The distributions of prediction errors for each tooth group were similar at T0 and T1, and no statistically significant differences were observed in the magnitude of errors(Fig. 8 ). This result suggests that the model’s accuracy did not significantly decrease after dental elements had undergone substantial movement due to orthodontic treatment, supporting, to some extent, its stability in predicting dental element width under viewpoint changes. Fig. 8. Open in a new tab Stability of Model Accuracy Across Tooth Position Changes: Error Distributions at Pre- (T0) and Post-treatment (T1) Stages. NS: no significant difference ( p > 0.05). Box plots show the distribution of prediction errors (versus 3D ground truth) for each tooth group at the initial (T0, blue) and post-treatment (T1, pink) stages. The stability of error distributions across time points indicates that the model’s estimation accuracy was not compromised by tooth movement Discussion The rapid advancement of digital dentistry has made it imperative to delegate highly repetitive and time-consuming anatomical measurement tasks to artificial intelligence (AI) to enhance clinical efficiency. Traditional tooth data acquisition relies on 3D intraoral scans or physical plaster models. The former carries the risk of permanent data loss in cases of device malfunction or software updates, while the latte r are often associated with data gaps during case reviews and treatment evaluations due to material fragility and storage limitations. Two-dimensional intraoral photographs, as a stable-format and low-cost archival medium [ 29 ], have long been regarded merely as visual references. In this study, we demonstrated that deep learning can estimate individual tooth widths from a single intraoral occlusal photograph, thus enabling the “salvage” recovery of anatomical parameters in cases of missing records. Furthermore, this strategy provides traceable data for secondary analyses such as the Bolton index and crowding assessment [ 1 ]. In forensic odontology and archaeological anthropology, this approach can also facilitate the rapid estimation of dentition dimensions using historical photographs or images of unearthed skulls, thereby facilitating identity verification and population migration studies. The proposed method extracts depth-informed features from 2D intraoral photos to predict tooth dimensions, offering low-cost, retrospective, and cross-scenario digital solution. The framework provides a transferable paradigm for anatomical quantification in 2D medical imaging, with potential extensions to tasks such as joint space evaluation or skin lesion area estimation. Our model enables fully automated and highly accurate tooth-width estimation from intraoral photographs. On an independent external test set ( n = 97 patients from a secondary center), it achieved an MAE of 0.34 mm and an ICC of 0.98, matching the established manual-measurement error range of 0.3–0.4 mm [ 30 , 31 ] and demonstrating reproducibility and stablity across diverse clinical settings. The computational efficiency was high, with a single jaw processed in approximately 0.075 s, highlighting the practical advantages of automated estimation. Notably, the AUC values obtained in this study were relatively lower compared to certain keypoint detection tasks, primarily due to the choice of normalization scale. In contrast to commonly used normalization units such as “interpupillary distance” or “face bounding box diagonal” in facial keypoint detection, the width of a single incisor was adopted as the normalization unit in this study. This approach is more stringent, as the incisor width is numerically smaller, causing the same pixel error to be amplified in normalized coordinates. Consequently, the PCK curve shifted downwards overall, resulting in a lower AUC. This strategy better aligns with the clinical demand for millimeter-level precision in dental practice. Therefore, the metrics obtained under this rigorous setting possess higher clinical relevance and interpretability. In previous similar studies, Ryu and colleagues successfully achieved the classification of tooth crowding severity and the diagnosis of orthodontic extraction [ 21 ]. However, their estimation of inter-tooth distances relied on two critical assumptions: [ 1 ] the use of average central incisor width in the Korean population as a reference standard, and [ 2 ] the presumption that the photographic angle was perfectly parallel to the occlusal plane. Furthermore, Hertig et al. quantified the degree of dental crowding in the anterior mandibular arch region [ 22 ], but similarly adopted an oversimplified assumption that pixel distance was directly proportional to actual physical distance, thereby neglecting depth information in spatial geometry. Nevertheless, since Hertig et al. focused exclusively on the mandibular anterior region, the impact of depth and photographic angle was relatively minimal. In fact, tooth spacing measurements based solely on 2D images are inherently subject to significant errors, as depth information is entirely lost during projection. This assumption leads to particularly pronounced inaccuracies when the shooting angle is not perfectly parallel to the occlusal plane or when dental arch rotation, crowding, or curvature exists. Teeth that appear adjacent in 2D projections may exhibit substantial spatial discrepancies in reality, thus resulting in cumulative deviations. Unlike traditional geometric measurement methods, our TDENet employs a regression paradigm. In this framework, the 2D keypoints are not treated as absolute physical scalars for distance calculation; instead, they serve as feature anchors to sample multi-level depth-aware representations from the local neighborhood. This design allows the model to perceive the high-dimensional spatial manifold, effectively mitigating the geometric ambiguity inherent in monocular 2D-to-3D inference. Guided by the principle of Consistency Regularization, our keypoint perturbation strategy (Fig. 4 ) encourages the model to generate invariant width outputs despite minor detection variations. By mapping shifted 2D coordinates to constant 3D ground-truth, the network is compelled to learn smooth and stable depth features, significantly enhancing the model’s robustness to varying camera poses and initial localization errors. Furthermore, our model functions as an implicit self-calibrator. By internalizing the anatomical morphological priors of over 12,000 teeth, the system decodes the absolute physical scale from purely visual and depth-informed cues, bypassing the need for a physical calibration marker. While 3D scanning remains the gold standard, this 2D-based approach provides a low-cost, high-efficiency alternative for large-scale screenings and teledentistry in resource-limited settings. Although overall performance was excellent, two systematic errors appeared in both cohorts.First, canine accuracy was consistently lower. R² values for maxillary and mandibular canines fell well below those of other teeth. The drop was sharpest in the mandibular canines, with internal R² at 0.38 and ICC at 0.73, and external R² at 0.38 and ICC at 0.55. Visual inspection showed that unfavourable camera angles or rotated canines often obscured the distal oblique ridge, making the distal contact point impossible to locate. This partial or complete occlusion affected 18.8% of internal images and 12.9% of external images. Its presence in both datasets confirms a universal imaging artefact rather than a centre-specific problem. Future studies should therefore add oblique lateral views and fuse them with standard frontal images to improve canine predictions.Second, R² values across all teeth were modest, with medians of 0.59 internally and 0.50 externally. This reflects the naturally narrow biological spread in tooth width; for example, 80% of maxillary central incisors lay between 8.12 mm and 8.85 mm. In such a low-variance setting, even small absolute errors markedly reduce R². Thus, MAE and ICC provide more meaningful measures of accuracy for this application [ 32 ]. The present study has the following limitations that need to be considered. First, regarding generalizability, although the model was externally validated across two independent clinical centers, both cohorts were exclusively derived from a Han Chinese population and primarily included permanent dentition. While the two centers utilized identical imaging equipment and followed the same standardized protocol to ensure high clinical reproducibility and stability, such demographic and technical similarity may limit the assessment of the model’s robustness against extreme variations in ethnicity, device types, or lighting conditions. Furthermore, as a model that learns statistical priors from data, its performance on anatomically extreme cases—such as teeth with congenital size anomalies (e.g., microdontia/macrodontia) that fall outside the primary distribution of our training set—has not been evaluated and represents a known boundary of the current approach. Meanwhile, The current model was developed and validated on cohorts dominated by permanent dentition. Its efficacy and accuracy in pediatric populations with mixed or primary (deciduous) dentition have not been established. The distinct morphology and developmental stages of teeth in these younger patients present a unique challenge that warrants dedicated training data and validation in future studies. Second, regarding imaging conditions, the model was developed and validated using high-quality, standardized occlusal photographs acquired under a strict imaging protocol. This controlled setting reflects the routine, standardized method for acquiring intraoral images in orthodontic practice, ensuring consistency for the initial proof-of-concept. However, its performance on non‑standard images—such as those taken from oblique angles, with varying camera‑to‑subject distances, under partial occlusion, or using different camera systems—has not been established. These factors may introduce variations in projective geometry and image resolution that could affect keypoint localization and subsequent width estimation. Therefore, the current workflow is recommended for use only with images that conform to the standardized occlusal acquisition protocol described in this work. Third, regarding clinical precision, although the model demonstrated satisfactory overall estimation performance, its accuracy in posterior tooth regions remains insufficient to meet high‑precision requirements for applications such as prosthetic rehabilitation and dental implantation. Looking ahead, future studies should build multi-ethnic aggregated datasets through international multi-center collaborations, incorporate more diverse population data and dental‑arch samples at different developmental stages to ensure broader demographic applicability. Conclusion In this study, we present a fully automated deep learning–based framework for estimating mesiodistal crown widths from standardize occlusal photographs. By integrating tooth keypoint detection with depth-informed feature representations and supervised regression, the proposed method enables accurate and efficient prediction of individual tooth dimensions without explicit camera calibration or physical scale references. The strong agreement between model-estimated values and clinically obtained reference measurements from 3D intraoral scans, together with consistent performance on an independent external dataset, supports the reliability of the proposed approach for retrospective research applications. Future studies will further investigate robustness across heterogeneous imaging devices and acquisition conditions. Supplementary Information Supplementary Material 1. (333.8KB, docx) Acknowledgements we would like to express our gratitude to EditSprings (https://www.editsprings.com) for the expert linguistic services provided. Abbreviations 3D Three-dimensional 2D Two-dimensional MAE Mean absolute error RMSE Root mean squared error R² The coefficient of determination CI Confidence interval AI Artificial intelligence CBAM-HRNet the Convolutional Block Attention Module-High-Resolution Network TDENet Tooth Distance Estimation Network HR-Net High-resolution network CBAM Convolutional block attention module MLP Multilayer perceptron CAGrad Conflict-Averse Gradient NME Normalized mean error PCK Percentage of correct keypoints AUC Area under the PCK curve ICC(3,1) Intraclass Correlation Coefficient for intra-rater reliability (single rater, absolute agreement) CV Cross-Validation Authors’ contributions L. T. Li and Y. L. Shi conceived and designed the study, performed data collection and analysis, and wrote the manuscript. D. Wang provided technical support, critically supervised the research, and revised the manuscript. W. X. Li critically supervised and revised the manuscript. T. T. Liu performed data collection and revised the manuscript. Y. Q. Ma revised the manuscript. M. Cao commented on and edited the manuscript, and served as the corresponding author. All authors reviewed the manuscript. Funding This work was supported by the New Technologies and Businesses Program, Hospital of Stomatology, Air Force Medical University [Grant No. LX-202015]; the China Postdoctoral Science Foundation [Grant No. 2024M764332]; and the Young Scientists Fund, State Key Laboratory of Oral & Maxillofacial Reconstruction and Regeneration [Grant No. 2024QN06]. The funders had no role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript. Data availability The datasets generated and/or analysed during the current study are available in the TDE repository, https://github.com/TT-WD/TDE . Declarations Ethics approval and consent to participate This study has been approved by the Medical Ethics Committee of the Third Affiliated Hospital of the Air Force Medical University (KQ-YJ-2025-127). All methods were carried out in accordance with relevant guidelines and regulations. Written consent was waived as the study used anonymized retrospective data. Consent for publication Not applicable. Competing interests The authors declare no competing interests. Footnotes Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Litong Li and Yulin Shi contributed equally to this work and share first authorship. References 1. Al-Tamimi T, Hashim HA. Bolton tooth-size ratio revisited. World J Orthod. 2005;6(3):289–95. [ PubMed ] [ Google Scholar ] 2. Alam MK, Shahid F, Purmal K, Sikder MA, Saifuddin M. Human Mesiodistal Tooth Width Measurements and Comparison with Dental Cast in a Bangladeshi Population. J Contemp Dent Pract. 2015;16(4):299–303. [ DOI ] [ PubMed ] [ Google Scholar ] 3. Othman SA, Harradine NWT. Tooth-size discrepancy and Bolton’s ratios: a literature review. J Orthod. 2006;33(1):45–51. discussion 29. [ DOI ] [ PubMed ] [ Google Scholar ] 4. Xie C, Sun M, He Z, Yu H. Digital intraoperative evaluation of restorative space and nontemplate-guided tooth preparation when replacing failed anterior restorations: a dental technique. J Prosthet Dent. 2024;S0022-3913(24):00279-8. [ DOI ] [ PubMed ] 5. Yang G, Chen Y, Li Q, Benítez D, Ramírez LM, Fuentes-Guajardo M, et al. Dental size variation in admixed Latin Americans: Effects of age, sex and genomic ancestry. PLoS ONE. 2023;18(5):e0285264. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Joda T, Brägger U. Patient-centered outcomes comparing digital and conventional implant impression procedures: a randomized crossover trial. Clin Oral Implants Res. 2016;27(12):e185–9. [ DOI ] [ PubMed ] [ Google Scholar ] 7. Emara A, Sharma N, Halbeisen FS, Msallem B, Thieringer FM. Comparative Evaluation of Digitization of Diagnostic Dental Cast (Plaster) Models Using Different Scanning Technologies. Dent J (Basel). 2020;8(3):79. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Christopoulou I, Kaklamanos EG, Makrygiannakis MA, Bitsanis I, Perlea P, Tsolakis AI. Intraoral Scanners in Orthodontics: A Critical Review. Int J Environ Res Public Health. 2022;19(3):1407. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Jackson TH, Kirk CJ, Phillips C, Koroluk LD. Diagnostic accuracy of intraoral photographic orthodontic records. J Esthet Restor Dent. 2019;31(1):64–71. [ DOI ] [ PubMed ] [ Google Scholar ] 10. Chang K, Beers AL, Bai HX, Brown JM, Ly KI, Li X, et al. Automatic assessment of glioma burden: a deep learning algorithm for fully automated volumetric and bidimensional measurement. Neuro Oncol. 2019;21(11):1412–22. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Wang J, Sun K, Cheng T, Jiang B, Deng C, Zhao Y, et al. Deep High-Resolution Representation Learning for Visual Recognition. IEEE Trans Pattern Anal Mach Intell. 2021;43(10):3349–64. [ DOI ] [ PubMed ] [ Google Scholar ] 12. Xu Z, Li B, Geng M, Yuan Y. AnchorFace. arXiv. 2021. Available from: http://arxiv.org/abs/2007.03221 . [cited 2025 May 15]. 13. Liskowski P, Krawiec K. Segmenting Retinal Blood Vessels With Deep Neural Networks. IEEE Trans Med Imaging. 2016;35(11):2369–80. [ DOI ] [ PubMed ] [ Google Scholar ] 14. Liu Y, Zuo S. Self-supervised monocular depth estimation for gastrointestinal endoscopy. Comput Methods Programs Biomed. 2023;238:107619. [ DOI ] [ PubMed ] [ Google Scholar ] 15. Angus LM, Mikołajczyk M, Cheung AS, Kasielska-Trojan AK. Validation of Breast Idea Volume Estimator Application in Transfeminine People. Plast Reconstr Surg Glob Open. 2024;12(9):e6131. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Ryu J, Lee YS, Mo SP, Lim K, Jung SK, Kim TW. Application of deep learning artificial intelligence technique to the classification of clinical orthodontic photos. BMC Oral Health. 2022;22(1):454. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Kuehnisch J, Meyer O, Hesenius M, Hickel R, Gruhn V. Caries Detection on Intraoral Images Using Artificial Intelligence. J Dent Res. 2022;101(2):158–65. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Yoon K, Jeong HM, Kim JW, Park JH, Choi J. AI-based dental caries and tooth number detection in intraoral photos: Model development and performance evaluation. J Dent. 2024;141:104821. [ DOI ] [ PubMed ] [ Google Scholar ] 19. You W, Hao A, Li S, Wang Y, Xia B. Deep learning-based dental plaque detection on primary teeth: a comparison with clinical assessments. BMC Oral Health. 2020;20(1):141. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Ragodos R, Wang T, Padilla C, Hecht JT, Poletta FA, Orioli IM, et al. Dental anomaly detection using intraoral photos via deep learning. Sci Rep. 2022;12(1):11577. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Ryu J, Kim YH, Kim TW, Jung SK. Evaluation of artificial intelligence model for crowding categorization and extraction diagnosis using intraoral photographs. Sci Rep. 2023;13(1):5177. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Hertig G, van Nistelrooij N, Schols J, Xi T, Vinayahalingam S, Patcas R. Quantitative tooth crowding analysis in occlusal intra-oral photographs using a convolutional neural network. Eur J Orthod. 2025;47(3):cjaf025. [ DOI ] [ PubMed ] [ Google Scholar ] 23. 《人工智能医疗器械注册审查指导原则》发布 - 中国知网. Available from: https://kns.cnki.net/kcms2/article/abstract?v=9oehDy4zW5ZrQs9qyFP9hxq9yRLrZCfLTWMAuDEVpW1LaFns6nJZ1WGI1hJPHtNW5ZvQkk2DCUX-_sOGwHQqKHG-05XVk5pSy657-FunmhLLk5QPprwZIRTya-jfyXCRDMIJ40z4LWiruzxKGRX1USs_qJqCndg3Yh23ffSfLn3mPPBuPw3fpA==&uniplatform=NZKPT&language=CHS . [cited 2025 Aug 18]. 24. Valle R, Buenaposada JM, Valdés A, Baumela L. Face alignment using a 3D deeply-initialized ensemble of regression trees. Comput Vis Image Underst. 2019;189:102846. [ Google Scholar ] 25. Pourramezan Fard A, Mahoor MH. Facial landmark points detection using knowledge distillation-based neural networks. Comput Vis Image Underst. 2022;215:103316. [ Google Scholar ] 26. Zhang J, Peng J, Wang K. Athlete posture estimation and analysis based on embodied artificial intelligence. Image Vis Comput. 2025;162:105598. [ Google Scholar ] 27. Jiang M, Tian Z, Yu C, Shi Y, Liu L, Peng T, et al. Intelligent 3D garment system of the human body based on deep spiking neural network. Virtual Real Intell Hardw. 2024;6(1):43–55. [ Google Scholar ] 28. Liu B, Liu X, Jin X, Stone P, Liu Q. Conflict-Averse Gradient Descent for Multi-task Learning. arXiv. 2024. Available from: http://arxiv.org/abs/2110.14048 . [cited 2025 Aug 19]. 29. Jin CX, Li MX, Yu H, Gao Y, Guo YP, Xia GS, et al. High-Fidelity 3D Imaging of Dental Scenes Using Gaussian Splatting. J Dent Res. 2025;104(9):964–72. [ DOI ] [ PubMed ] [ Google Scholar ] 30. Santoro M, Galkin S, Teredesai M, Nicolay OF, Cangialosi TJ. Comparison of measurements made on digital and plaster models. Am J Orthod Dentofac Orthop. 2003;124(1):101–5. [ DOI ] [ PubMed ] [ Google Scholar ] 31. Asquith J, Gillgrass T, Mossey P. Three-dimensional imaging of orthodontic models: a pilot study. Eur J Orthod. 2007;29(5):517–22. [ DOI ] [ PubMed ] [ Google Scholar ] 32. Deng W, Liu Y, Hu J, Guo J. The small sample size problem of ICA: A comparative study and analysis. Pattern Recogn. 2012;45(12):4438–50. [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary Material 1. (333.8KB, docx) Data Availability Statement The datasets generated and/or analysed during the current study are available in the TDE repository, https://github.com/TT-WD/TDE . Articles from Head & Face Medicine are provided here courtesy of BMC ACTIONS View on publisher site PDF (2.1 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 15045 · SHA-256 0d8dd46020d42cd4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.