The Challenge of Diagnosing Pediatric Laryngeal Clefts
Pediatric laryngeal clefts, congenital abnormalities where the tissue between the vocal cords does not fully close, present a significant diagnostic challenge. Historically, descriptions of laryngeal anatomy, particularly the interarytenoid region, have varied widely among clinicians and institutions. This ambiguity has led to discrepancies in identifying a normal larynx, a deep interarytenoid notch (DIN), and a Benjamin-Inglis type I laryngeal cleft (T1C). The International Pediatric Otolaryngology Group (IPOG) highlighted this issue in their 2017 consensus guidelines for T1Cs, noting that many members relied solely on visual inspection for diagnosing DINs, while others employed specific height measurements without a standardized approach.
The critical importance of precise differentiation lies in its direct impact on patient management. Accurate identification of the severity of a laryngeal cleft, and its associated functional sequelae such as breathing difficulties and feeding problems, dictates whether a child may require surgical intervention. Without a unified assessment tool, the risk of misdiagnosis or delayed diagnosis increases, potentially compromising a child’s health and development.
Development and Validation of the IAAP
In response to this diagnostic inconsistency, the Interarytenoid Assessment Protocol (IAAP) was developed to provide a standardized method for describing the interarytenoid mucosal height (IAMH). The protocol involves a systematic approach using laryngoscopy and specialized instruments. It begins with placing the patient in suspension, followed by the insertion of laryngeal distending forceps to optimize visualization. Clinicians then palpate the interarytenoid musculature for subjective qualities before employing a right-angle microlaryngeal probe. This probe is carefully maneuvered to determine the point of initial mucosal contact, which is then designated as the IAMH. The IAMH is subsequently categorized into one of five levels relative to key laryngeal landmarks: the false vocal folds (FVFs), ventricle, and true vocal folds (TVFs).
A previous single-institution validation study had shown promising results regarding the IAAP’s inter- and intra-rater reliability. However, the need for a broader, multi-institutional validation was paramount to confirm its generalizability and clinical utility across diverse healthcare settings. This latest research project aimed to fulfill that need by conducting a comprehensive multi-institutional validation study to rigorously assess the IAAP’s reliability in evaluating the IAMH.
Methodology: A Multi-Institutional Blinded Review
The study employed a retrospective review of 30 endoscopic videos documenting IAAP procedures from the Seattle Children’s Hospital Otolaryngology Database. These procedures were performed by either attending pediatric otolaryngologists or supervised residents and fellows at that institution. Institutional Review Board (IRB) approval was secured from Seattle Children’s Hospital.

Sample Selection and Video Review:
A total of 115 records were initially reviewed for inclusion. The selection criteria required videos to be of sufficient quality for clear visualization of the IAAP procedure. Videos were excluded if they exhibited inadequate visualization of the protocol’s steps (32 videos), poor overall quality (8 videos), or if the target number of recordings for a specific IAMH category had already been met (45 videos). This meticulous screening process resulted in a final sample of 30 de-identified endoscopic videos.
Blinded Multi-Institutional Assessment:
Ten fellowship-trained pediatric otolaryngologists from 10 different institutions participated in the reliability assessment. Each rater was tasked with evaluating the 30 videos. Crucially, the raters were blinded to any identifying information of the patients and the number of videos representing each IAMH category. They were provided with the IAAP Scoring Sheet, which included illustrative examples of the right-angle probe’s position at various IAMH levels.
The raters were instructed to focus solely on assessing the anatomy based on the contact point of the microlaryngeal probe as demonstrated by the surgeon in the video, explicitly avoiding any independent assessment of surgical technique. To further ensure objectivity and prevent bias, the videos were edited to exclusively showcase the IAAP procedures, omitting any additional surgical steps. The order in which the videos were presented to the raters was randomized for both assessment periods. To gauge intra-rater reliability, each rater reviewed the same set of videos twice, with a three-month interval between the assessments. This allowed for a robust evaluation of consistency in individual rater performance over time.
Statistical Analysis:
The study calculated intra-class correlation (ICC) coefficients using two-way models to assess both inter-rater (agreement between different raters) and intra-rater (agreement within the same rater over time) reliability. The degree of agreement was categorized using established benchmarks: slight (0.00–0.20), fair (0.21–0.40), moderate (0.41–0.60), strong (0.61–0.80), and near complete agreement (0.81–1.00). Statistical analyses were performed using R Project for Statistical Computing, with a significance level set at P < 0.05.
Results: Strong Reliability Demonstrated
The study yielded compelling results, indicating a high degree of reliability for the IAAP when applied across multiple institutions. The cohort of 30 patients whose videos were analyzed comprised a median age of 4.9 years, with a range spanning from one month to 20 years. The gender distribution was 30% female and 70% male.
Inter-Rater Reliability:
The inter-rater reliability, measuring the consistency of assessments among different clinicians, demonstrated strong agreement. On the first assessment round, the inter-rater reliability was calculated at 0.74 (95% confidence interval [CI] 0.63 to 0.84). In the second assessment round, conducted three months later, the inter-rater reliability remained robust at 0.75 (95% CI 0.63 to 0.85). These figures fall within the "strong" to "near complete agreement" range, signifying that different clinicians are highly likely to arrive at similar conclusions when evaluating the same IAAP video.
Intra-Rater Reliability:
The intra-rater reliability, assessing the consistency of a single rater’s assessments over time, also showed significant strength. Overall intra-rater reliability ranged from 0.49 to 0.89, with an overall test-retest reliability of 0.75 (95% CI 0.69 to 0.79). This indicates that individual raters were highly consistent in their own evaluations when reassessing the videos after a three-month interval.

Levels of Agreement:
Further analysis revealed specific insights into the distribution of agreement. In 7 instances (23%) out of the total 500 separate reviews, there was 100% agreement among the 10 participants on the IAAP Classification. These unanimous classifications included "above FVF," "at FVF," "in ventricle," and "at TVF." More broadly, in nearly half of the cases (46.6%), raters assigned IAAP classifications that were within one level of each other, indicating a high degree of consensus. While there were instances of greater spread in classifications (two or more levels of disagreement in 30% of cases), the overall trend pointed towards substantial agreement.
Implications and Future Directions
The findings of this multi-institutional validation study carry significant implications for the field of pediatric otolaryngology. The IAAP’s demonstrated strong inter- and intra-rater reliability in visual assessment provides a much-needed standardized tool for evaluating the interarytenoid mucosal height. This standardization is crucial for accurately diagnosing laryngeal clefts, particularly the subtle distinctions between normal anatomy, deep interarytenoid notches, and Type I laryngeal clefts.
Enhancing Diagnostic Accuracy:
By offering a consistent method for describing IAMH, the IAAP can help mitigate the diagnostic variability that has historically plagued the assessment of laryngeal clefts. This improved diagnostic accuracy can lead to more timely and appropriate treatment decisions, potentially reducing the need for invasive diagnostic procedures and ensuring that children receive the necessary surgical or non-surgical interventions without delay.
Foundation for Outcome Studies:
The reliability of the IAAP lays the groundwork for more robust outcome studies in pediatric airway disorders and pharyngeal dysphagia. With a standardized assessment of laryngeal anatomy, researchers can more accurately correlate specific anatomical findings with functional outcomes, leading to a deeper understanding of the disease processes and the effectiveness of various interventions. This could pave the way for evidence-based guidelines that are more precise and universally applicable.
Acknowledging Limitations and Future Research:
The authors acknowledge that the IAAP’s overall effectiveness is also influenced by the precision of surgical technique. While this study focused on the reliability of visual interpretation alone through pictorial analysis, future research is warranted to explore how variations in surgical technique might impact the protocol’s overall clinical utility. Investigating the correlation between IAAP classifications and objective measures of airway function, such as aerodynamic assessments or dynamic imaging, would further enhance the protocol’s value.
Furthermore, continued efforts to educate and train pediatric otolaryngologists on the IAAP’s standardized methodology will be essential for its widespread adoption and consistent application. As the understanding and management of pediatric laryngeal clefts evolve, the IAAP stands as a significant advancement, promising to bring greater clarity and consistency to a critical area of pediatric airway care. This research marks a vital step toward ensuring that all children with laryngeal clefts receive the most accurate diagnosis and the most effective care possible, regardless of where they are treated.
