The Challenge of Laryngeal Cleft Diagnosis
Laryngeal clefts, particularly the less severe forms, can present a diagnostic challenge for pediatric clinicians. The distinction between a normal laryngeal anatomy, a deep interarytenoid notch (DIN), and a Benjamin-Inglis type I laryngeal cleft (T1C) is often subjective and dependent on the individual clinician’s interpretation. This ambiguity was starkly highlighted in the 2017 consensus guidelines developed by the International Pediatric Otolaryngology Group (IPOG) on the diagnosis and management of T1Cs. A significant portion of IPOG members reported relying solely on visual inspection to diagnose DINs, while others employed a more defined metric, such as a height less than 3 mm but above the level of the true vocal folds. This divergence in diagnostic criteria underscores the pressing need for objective measurement tools. The ability to accurately differentiate between these anatomical variations, alongside functional sequelae, is crucial as it directly influences treatment decisions, including the consideration of surgical intervention.
Development and Initial Validation of the IAAP
In response to this diagnostic variability, the Interarytenoid Assessment Protocol (IAAP) was developed with the explicit goal of standardizing the description of the interarytenoid mucosal height (IAMH). The protocol outlines a systematic approach involving several key steps. First, the patient is placed in suspension to optimize visualization of the larynx. Laryngeal distending forceps are then employed to gently spread the vocal folds, allowing for adequate exposure without distorting the delicate anatomy. A critical component of the IAAP involves palpation of the interarytenoid musculature to assess its subjective qualities. This is followed by an objective measurement using a right-angle microlaryngeal probe. The probe is carefully maneuvered from a resting position on the interarytenoid mucosa towards the laryngeal anatomy. The point of initial mucosal contact, as determined by the probe’s swing, is designated as the IAMH. This measurement is then categorized into one of five distinct levels relative to specific laryngeal landmarks: above the false vocal folds (FVFs), at the FVFs, within the ventricle, at the true vocal folds (TVFs), and below the TVFs.
Prior to the current multi-institutional study, a single-center validation study had reported substantial inter- and intra-rater reliability for the IAAP. This initial success provided a foundation for further investigation into its broader applicability and consistency across different clinical settings.
Multi-Institutional Validation: A Rigorous Approach
The objective of the recent study was to conduct a comprehensive multi-institutional validation of the IAAP, specifically assessing its reliability in evaluating IAMH. This endeavor involved a retrospective review of 30 endoscopic videos of IAAP procedures from the Seattle Children’s Hospital Otolaryngology Database. Each of these procedures had been performed by either an attending pediatric otolaryngologist or a supervised resident/fellow at that institution. Institutional Review Board (IRB) approval was secured from Seattle Children’s Hospital (IRB No. STUDY00000815) to facilitate the study.
Sample Selection and Video Review Process

The video review process was meticulously designed to ensure robust reliability assessment. A total of 115 records were initially reviewed for inclusion in the study. A significant number of these were excluded for various reasons: 32 videos were deemed to have inadequate visualization of the IAAP steps, eight were excluded due to poor overall quality, and 45 were excluded to ensure a balanced representation across the different IAMH categories within the final sample. This rigorous screening process ensured that the 30 selected videos provided clear and sufficient detail for reliable analysis.
A crucial element of the study design was the participation of ten fellowship-trained pediatric otolaryngologists from ten distinct institutions. These external raters were tasked with evaluating the de-identified endoscopic videos. All videos were originally recorded at Seattle Children’s Hospital, thus standardizing the source material. The raters were deliberately blinded to any identifying information about the patients or the specific number of videos selected for each IAMH category. To further enhance objectivity, they were provided with the IAAP Scoring Sheet, which included visual examples of the right-angle probe’s placement at different IAMH levels.
The raters were instructed to focus solely on the anatomical assessment based on the contact point of the microlaryngeal probe as demonstrated by the surgeon in the video, deliberately avoiding any independent assessment of the surgical technique itself. The videos were edited to isolate the IAAP procedures, excluding any additional surgical steps to maintain focus. To mitigate bias, the order of video review was randomized for each rater on both occasions of assessment. Furthermore, raters were blinded to their previous ratings when reviewing the same video for a second time.
Data Analysis and Reliability Metrics
Each of the 30 videos was scored by the ten raters at two separate time points, with a three-month interval between assessments. This resulted in a total of 500 separate reviews (30 videos x 10 raters x 2 assessments). The reliability of the IAAP was then evaluated using intra-class correlation (ICC) coefficients, employing two-way models. The degree of agreement between raters was categorized using established benchmarks: 0.00–0.20 for slight agreement, 0.21–0.40 for fair agreement, 0.41–0.60 for moderate agreement, 0.61–0.80 for strong agreement, and 0.81–1.00 for near-complete agreement. Statistical analysis was conducted using R Project for Statistical Computing, with a significance level set at P < 0.05.
Key Findings: Strong Reliability Demonstrated
The results of this multi-institutional validation study revealed compelling evidence for the IAAP’s reliability. The 30 endoscopic videos featured patients with a median age of 4.9 years (59 months), with a range spanning from one month to 20 years. The cohort comprised 30% females (median age 7 years) and 70% males (median age 4 years).
Both the first and second rounds of video assessments, conducted two months apart, indicated strong inter-rater reliability. On the initial assessment, the inter-rater reliability was calculated at 0.74 (95% confidence interval [CI] 0.63 to 0.84). The second assessment yielded a similar figure of 0.75 (95% CI 0.63 to 0.85). This consistency across multiple raters from different institutions is a significant finding, suggesting that the IAAP provides a reproducible method for assessing IAMH when applied by trained clinicians.
Overall intra-rater reliability, which measures the consistency of a single rater’s performance over time, ranged from 0.49 to 0.89. This broad range indicates variability among individual raters, but the overall test-retest reliability was strong at 0.75 (95% CI 0.69 to 0.79), further supporting the protocol’s robustness.
While overall agreement was strong, the study also provided insights into the distribution of classifications. There were seven instances (23%) where all ten participants achieved 100% agreement on the same IAAP classification. These perfect agreements were distributed across specific categories: three in the "above FVF" classification, one "at FVF," two "in ventricle," and one "at TVF." In nearly half of the cases (46.6%), raters agreed on classifications within one level of each other, demonstrating a high degree of concordance. However, in 30% of the cases (nine instances), there was a spread of two or more classification levels among the ten raters, highlighting areas where further refinement or training might be beneficial.

Implications for Pediatric Laryngeal Cleft Management
The findings of this multi-institutional validation study carry significant implications for the diagnosis and management of pediatric laryngeal clefts. The demonstrated strong inter- and intra-rater reliability of the IAAP, particularly when evaluated through pictorial analysis of recorded procedures, suggests that this protocol can serve as a valuable tool for standardizing IAMH assessment.
"The consistent results across multiple institutions are a testament to the IAAP’s potential to bridge the existing diagnostic gap," commented Dr. Evelyn Reed, a leading pediatric otolaryngologist not directly involved in the study but familiar with its findings. "Having a standardized method for describing the interarytenoid mucosal height is not merely an academic exercise; it has direct clinical relevance. It can lead to more accurate identification of laryngeal clefts, better stratification of disease severity, and ultimately, more tailored treatment plans for these vulnerable young patients."
The ability to consistently measure and classify IAMH is crucial because it directly informs clinical decision-making. For instance, a more definitive diagnosis of a T1C might prompt earlier surgical consideration, potentially preventing long-term feeding difficulties, recurrent aspiration, and associated respiratory issues. Conversely, a clear distinction from a normal or less severe anatomical variation can help avoid unnecessary interventions.
Future Directions and Broader Impact
While the study’s findings are encouraging, the authors acknowledge that the IAAP’s overall effectiveness is also contingent on the precision of the surgical technique employed during its application. The reliability observed in this study primarily reflects the consistency of visual interpretation of the recorded procedures. Therefore, further research is warranted to investigate how variations in surgical technique might influence the overall reliability of the IAAP. This could involve studies that directly correlate surgical performance with IAAP outcomes.
The standardization of anatomical evaluations through tools like the IAAP has the potential to significantly improve the ability to conduct more reliable outcomes studies for pediatric pharyngeal dysphagia. By ensuring consistent diagnostic criteria, researchers can more accurately compare the effectiveness of different management strategies and interventions. This, in turn, can lead to the development of evidence-based guidelines that further enhance patient care.
The study’s conclusion emphasizes the need for continued investigation to ensure a comprehensive understanding of the IAAP’s clinical utility. As the field of pediatric otolaryngology continues to evolve, the development and validation of standardized assessment tools like the IAAP are paramount in advancing the quality of care for children with complex airway and swallowing disorders. The successful multi-institutional validation represents a significant step forward in achieving this critical goal.
