To address this critical need for standardization, the Interarytenoid Assessment Protocol (IAAP) was developed. This protocol aims to provide a consistent and objective method for describing the interarytenoid mucosal height (IAMH). The IAAP procedure involves a multi-step process: suspending the patient’s larynx to optimize visualization, employing laryngeal distending forceps to gently expand the laryngeal structures, and then palpating the interarytenoid musculature to assess its subjective qualities. Following this, a right-angle microlaryngeal probe is carefully maneuvered to determine the precise point of initial mucosal contact as it sweeps from the anterolateral interarytenoid mucosa towards the laryngeal anatomy. This point of contact defines the IAMH, which is then categorized into one of five distinct levels relative to key laryngeal landmarks: the false vocal folds, the ventricle, and the true vocal folds. While a previous single-institution validation study indicated substantial inter- and intra-rater reliability for the IAAP, the medical community has awaited broader validation across multiple institutions to confirm its robustness and generalizability. The current study endeavors to fill this gap by undertaking a multi-institutional validation of the IAAP, specifically evaluating its reliability in the assessment of IAMH.
Background: The Challenge of Laryngeal Cleft Diagnosis
Laryngeal clefts are congenital anomalies characterized by an incomplete fusion of the posterior larynx. The severity of these clefts can range from a subtle indentation (deep interarytenoid notch) to a more significant gap extending into the vocal cords or beyond. These conditions can lead to a spectrum of symptoms, including feeding difficulties, aspiration, recurrent pneumonia, and voice abnormalities. The diagnosis and subsequent management of these subtle anatomical variations have historically posed a significant challenge for pediatric otolaryngologists.
The lack of a universally accepted definition for what constitutes a "normal" interarytenoid space versus a pathological cleft has led to considerable variation in diagnostic practices. This variability can result in delayed or inappropriate treatment for affected children. For instance, a child with a significant laryngeal cleft might be misdiagnosed as having a simple vocal cord dysfunction, while another with a minor anatomical variation might be subjected to unnecessary surgical intervention. The 2017 IPOG guidelines highlighted this issue, revealing that many experts relied on subjective visual assessment to identify DINs, leading to a broad range of interpretations. The proposed quantitative measure of less than 3mm height for a DIN, situated above the true vocal folds, represented an attempt to introduce more objectivity, but the practical application and validation of such criteria remained a crucial next step.
The IAAP protocol emerged as a response to this diagnostic dilemma. By standardizing the measurement of IAMH, it aims to provide clinicians with a more reliable tool to differentiate between normal anatomy and varying degrees of laryngeal clefts. This, in turn, is expected to lead to more consistent patient stratification and informed treatment decisions, particularly regarding surgical intervention, which is typically reserved for more severe cases or those with significant functional impairment.
Methodology: A Rigorous Multi-Institutional Validation
To rigorously assess the reliability of the IAAP, a retrospective review was conducted utilizing 30 endoscopic videos from the Seattle Children’s Hospital Otolaryngology Database. Each of these video recordings documented IAAP procedures performed by either experienced attending pediatric otolaryngologists or residents and fellows under the direct supervision of an attending physician at that institution. The study received Institutional Review Board (IRB) approval from Seattle Children’s Hospital (IRB No. STUDY00000815), ensuring ethical compliance.

Sample Selection and Video Review
The selection process for the videos was meticulous. A total of 115 records were initially reviewed for inclusion. From this pool, 32 videos were excluded due to insufficient visualization of the critical steps of the IAAP procedure. An additional eight videos were disqualified due to poor overall video quality, which would have compromised accurate assessment. Furthermore, 45 videos were excluded because the target number of cases for specific IAMH categories had already been met, ensuring a balanced representation across the potential classifications. This left a final sample of 30 high-quality videos for the reliability analysis.
The core of the validation involved a blinded review process by independent experts. Ten fellowship-trained pediatric otolaryngologists, each from a distinct institution, participated in the evaluation of the IAAP. This multi-institutional approach was crucial for assessing the protocol’s generalizability beyond a single clinical setting. Each of these ten raters independently evaluated the same 30 de-identified endoscopic videos. Importantly, all videos were originally recorded at Seattle Children’s Hospital, but the raters were external to this institution, minimizing institutional bias.
To ensure objectivity, the raters were blinded to any identifying information of the patients and the specific number of videos assigned to each IAMH category. They were provided with the standardized IAAP Scoring Sheet, which included illustrative examples of the right-angle probe’s positioning at various IAMH levels. The raters were specifically instructed to focus solely on the anatomical depiction of the probe’s contact point as demonstrated by the surgeon in the video, refraining from any independent assessment of the surgical technique itself. To further standardize the review, the videos were edited to exclusively feature the IAAP procedure, excluding any ancillary or concurrent surgical interventions.
The reliability assessment was conducted over two separate time points, with an interval of three months between each viewing. This temporal separation was designed to evaluate both inter-rater reliability (agreement among different raters) and intra-rater reliability (consistency of a single rater over time). The order in which the videos were presented to the raters was randomized on both occasions. Crucially, when reviewing the videos for the second time, raters were not provided with their previous assessment of the same video, thus preventing any recall bias. This rigorous methodology resulted in a total of 500 separate video reviews (30 videos x 10 raters x 2 time points), providing a substantial dataset for statistical analysis.
Data Analysis
The statistical analysis focused on quantifying the degree of agreement among the raters. Intra-class correlation (ICC) coefficients, employing two-way random-effects models, were computed to assess both inter- and intra-rater reliability. The interpretation of the ICC values followed established guidelines:
- 0.00–0.20: Slight agreement
- 0.21–0.40: Fair agreement
- 0.41–0.60: Moderate agreement
- 0.61–0.80: Strong agreement
- 0.81–1.00: Near complete agreement
Statistical analyses were performed using R Project for Statistical Computing, with a significance level set at P < 0.05.
Results: Demonstrating Strong Reliability
The study’s findings indicate a promising level of reliability for the IAAP protocol, even when applied across multiple institutions and by different evaluators. The 30 endoscopic videos included patients with a median age of 4.9 years (interquartile range: 59 months), with ages spanning from one month to 20 years. The cohort comprised 30% females (median age of 7 years) and 70% males (median age of 4 years).

The primary analysis of inter-rater reliability revealed strong agreement between the ten participating fellowship-trained pediatric otolaryngologists. On the first video assessment (Part 1), the inter-rater reliability was calculated at 0.74 (95% confidence interval [CI] 0.63 to 0.84). This strong agreement persisted on the second video assessment (Part 2), with an inter-rater reliability of 0.75 (95% CI 0.63 to 0.85). These figures fall within the "strong agreement" category, suggesting that the IAAP provides a consistent framework for experienced clinicians to interpret IAMH.
Furthermore, the intra-rater reliability, which measures the consistency of individual raters over time, also demonstrated robust results. Overall intra-rater reliability ranged from 0.49 to 0.89, indicating moderate to near-complete agreement for individual raters when reassessing the same videos after a three-month interval. The overall test-retest reliability for the IAAP was 0.75 (95% CI 0.69 to 0.79), further supporting the protocol’s consistency.
Digging deeper into the agreement rates, the study found that in 7 instances (23% of the total assessments), there was 100% agreement among all 10 participants regarding the IAAP Classification. These instances of perfect consensus included three classifications of "above FVF," one "at FVF," two "in ventricle," and one "at TVF." While perfect agreement was not achieved in all cases, a significant portion of the assessments showed a high degree of concordance. In nearly half of the evaluations (14 instances, or 46.6%), raters assigned IAAP classifications that were within one level of each other, demonstrating a high level of concordance even when not in complete agreement. However, in 9 cases (30% of the evaluations), there was a spread of two or more classification levels among the 10 raters, highlighting areas where further refinement or additional training might be beneficial.
Implications and Future Directions
The findings of this multi-institutional validation study are highly encouraging, suggesting that the Interarytenoid Assessment Protocol (IAAP) offers a reliable method for assessing the interarytenoid mucosal height (IAMH) when evaluated through video analysis. The observed strong inter- and intra-rater reliability indicates that the protocol can consistently guide clinicians in categorizing laryngeal anatomy, a crucial step in the diagnosis and management of pediatric laryngeal clefts. This standardization holds the potential to significantly reduce diagnostic variability that has historically plagued this field.
However, the study authors cautiously note that the reliability demonstrated is primarily a reflection of visual interpretation. The effectiveness of the IAAP is inherently linked to the precision with which the surgical technique is performed. Variations in how the laryngoscope is suspended, how the laryngeal distending forceps are applied, and the subtle manipulation of the microlaryngeal probe can all influence the visual information available to the rater. Therefore, while the protocol itself appears robust for visual assessment, further research is warranted to explore how potential variations in surgical technique might impact the overall reliability and clinical utility of the IAAP. Understanding these surgical nuances will be critical for a comprehensive assessment of the protocol’s real-world applicability.
The implications of a standardized and reliable diagnostic tool like the IAAP extend beyond individual patient care. Improved diagnostic accuracy can lead to more precise patient stratification, enabling clinicians to better identify those who would benefit most from surgical intervention. This, in turn, can optimize resource allocation and potentially improve surgical outcomes. Moreover, a standardized approach facilitates more reliable outcome studies. By ensuring that patients are classified consistently, researchers can conduct more robust studies to evaluate the effectiveness of different treatment modalities and surgical techniques for pediatric laryngeal clefts. This enhanced ability to perform reliable outcome studies is vital for advancing the field of pediatric otolaryngology and improving the management of conditions such as pediatric pharyngeal dysphagia, which can be a consequence of laryngeal anomalies.
Looking ahead, the IAAP protocol represents a significant step towards achieving a more objective and consistent approach to laryngeal cleft assessment. Future research should focus on prospective validation studies, incorporating real-time clinical application and exploring the impact of surgical technique on IAAP reliability. Continued efforts in standardization and validation will undoubtedly contribute to more accurate diagnoses, improved patient management, and ultimately, better outcomes for children with laryngeal clefts. The collaborative nature of this study, involving multiple institutions and experienced clinicians, underscores the commitment of the pediatric otolaryngology community to advancing patient care through evidence-based practice and methodological refinement.
