Activity 1: Evaluating AI Use-Cases
Background
Autoverification refers to the automated release of laboratory test results that are unlikely to have been the result of a preanalytical or analytical error, bypassing manual review by a technologist or pathologist. The primary objective is to improve laboratory efficiency, reduce turnaround time, and minimize the risk of human error in the post-analytical phase. Traditional autoverification rules are typically built via LIS rule sets that incorporates analytic flags, delta checks, reference intervals, and other decision thresholds derived from expert consensus or regulatory guidance (e.g., CLSI AUTO10-A).
However, conventional autoverification systems can struggle with the complexity and nuance of real-world patient data, particularly in the presence of atypical clinical scenarios or concurrent pre-analytical and analytical confounders. Errors may occur due to rigid or incomplete rules, leading either to inappropriate autoverification or excessive manual holds.
AI–driven autoverification aims to enhance performance by learning patterns from large datasets of historical results, including rare or complex scenarios. By integrating multiple data sources (patient demographics, clinical context, analyzer flags, historical trends), AI-based autoverification systems have the potential to optimize both sensitivity (error detection) and specificity (avoiding unnecessary manual review), improving both quality and operational efficiency.
Background
Serum protein electrophoresis (SPEP) is a critical diagnostic tool for identifying and characterizing monoclonal gammopathies and other disorders of serum protein distribution. Accurate interpretation of SPEP patterns requires specialized expertise to distinguish between polyclonal and monoclonal patterns, recognize atypical presentations, and correlate findings with clinical and laboratory context. Errors in interpretation can result from subtle pattern recognition challenges, background artifacts, or the presence of confounding clinical conditions (e.g., acute-phase response, nephrotic syndrome).
AI-based SPEP interpretation could automate the identification of monoclonal peaks and other clinically significant patterns. By training on large sets of expert-annotated electrophoresis tracings, such systems can support standardized, accurate, and efficient result interpretation.
Background
Clinical laboratories often provide interpretive comments for complex or high-stakes test results, such as endocrine testing, toxicology, coagulation studies, protein electrophoresis, or molecular diagnostics. These comments may synthesize current and prior laboratory results, relevant clinical context, specimen quality issues, medication effects, and recommended follow-up testing. Drafting them can be time-consuming and requires careful attention to accuracy, tone, scope of practice, and institutional reporting standards.
A generative AI tool could draft preliminary interpretive comments for laboratory professionals to review, edit, and approve before release. By combining structured laboratory data with a controlled prompt template and approved reference content, such a system could improve consistency and reduce documentation burden. However, it could also introduce clinically plausible but incorrect statements, overstate certainty, omit important caveats, or generate recommendations that do not match local practice.
Task
Discuss the questions below within your groups. Each tab focuses on a unique step in the complete machine learning life cycle, with its own objectives and considerations. A good AI use-case will have clear, justifiable answers to each of these questions.
1) What should the model be predicting?
2) What should we use as a ground truth label?
3) What should be included in the training data set?
1) How should performance be measured?
2) What thresholds should be considered acceptable?
3) What data should be used in the validation study?
1) What potential benefits could be expected?
2) What potential harms might be introduced?
3) Will the predictions be helpful in changing practice?
1) How will we monitor performance?
2) What will we do if performance deteriorates?
1) Are there human factors to consider?
2) Are there regulatory requirements to consider?the