CQMs in the age of LLMs: Increasing flexibility and reducing burden

The Reality Behind "e"CQMs

Electronic Clinical Quality Measures (eCQMs) are central to Value-Based Care, aiming to automate quality assessment from EHR data. However, the "electronic" aspect often masks significant manual effort and system limitations. This exploration delves into these challenges and proposes a more intelligent, flexible future for quality measurement, especially in the age of Large Language Models (LLMs).

The Promise vs. The Lived Reality

The Promise:

  • Auto-calculation from EHR data.
  • No extra clicks for clinicians.
  • Seamlessly "works on EHR system."
  • Accurate reflection of care.

The Lived Reality:

  • Significant manual workflow for clinicians & admins.
  • "Smuggled" clinician/admin time to meet specs.
  • Data often reflects "checklist compliance" not just care.
  • Standardized EHR data outputs often misalign with specific eCQM needs, requiring manual reconciliation.

Example: Vendor "Checklist" Snippet

[ ] PHQ-9 Score Entered in Form_XYZ
[ ] Follow-up Code_ABC Added to Plan Tab

The Hidden Burden: Evidence from EHR Manuals

Analysis of guidance from over 20 EHRs (2020-2024) reveals recurring pain points. These aren’t edge cases—every major EHR requires extra steps beyond normal charting.

Structured Fields or It Doesn’t Count: Answers must land in discrete flowsheet cells or specific template fields, not just free-text.
Separate, Coded Follow-Up Entry: A positive screen needs a *second* distinct, coded entry for the follow-up plan (e.g., referral, medication).
Extra Quality Checklist/Tab: Many EHRs use separate "quality modules" or "checklists" outside the main clinical note, forcing double documentation.
Local Build & Mapping Upkeep: Significant IT work needed to enable templates, alerts, and ensure data fields map correctly to eCQM logic, especially after upgrades.

The Data Integrity Gap: Impact on Population Analytics

This reliance on specific, manual workflows undermines the reliability of eCQM data for true population health analytics. The effort distribution further highlights this challenge.

  • "Aggregate & compute" fails if all data sources didn't follow the exact "special prep steps."
  • Cannot reliably "turn on" a new measure retroactively if specific documentation wasn't done.
  • Distorts "report-only" pilots: Manual efforts during pilots make measures look easier to implement than they are in sustained, real-world practice.

Effort Allocation: Improving Quality vs. "Making the Measure Happy"

Improving Actual Patient Care Quality ~30%
"Making the Measure Happy" ~70%

(EHR config, extra clicks, data validation, admin tasks for compliance, make work, etc.)

(Illustrative effort distribution based on common eCQM reporting challenges)

Case Study: The CMS2 (Depression Screening) Gauntlet

The PHQ-9, a common depression questionnaire, often highlights the eCQM disconnect. Clinically appropriate care may occur, yet denominator/numerator credit hinges on multiple manual configurations and data linkage steps.

(A more realistic, burdensome workflow)

Initial EHR Setup (Manual Admin Task):
Configure specific PHQ-9 form & associate with LOINC code.
⬇️
📱 App Check-in / Portal:
Patient completes PHQ-9 form.
⬇️
🖱️ EHR - Clinician Task:
Link completed form to current encounter + new diagnosis (if applicable).
⬇️
📝 QM Portal (Post-Encounter Staff Task):
Review encounter hx; Manually record/link follow-up as distinct "intervention".
⬇️
💡 "e"CQM Calculated:
Only after all prior manual configuration & data linkage steps are complete.

A Foundational Principle: Capture Once, Compute Anywhere

Let EHRs focus on comprehensive care documentation. Let a separate, intelligent service compute any measure—today or retrospectively.

EHR SystemsFocus: Rich Care Documentation (Notes, Orders, Problems, etc.)
decoupled from
Measure Computation LayerFocus: Applying Logic to ALL Available Data

Blueprint for Change: A Flexible, Intelligent Architecture

This proposed architecture can be built once and, in principle, deployed against any certified EHR to overcome current limitations.

Comprehensive Data Ingestion & Interaction
Bulk FHIR Export: FHIR APIs (Structured) + Full Clinical Notes (Narrative)
Full EHI Export (Backstop): All Structured Data (incl. Non-Standardized for LLM use)
UI Automation Agents (Backstop): For Non-API Accessible Data/Actions
⬇️
Intelligent Data Structuring & Enrichment
LLM Agents: Use structured & unstructured data to classify patients according to CQM criteria.
⬇️
Measure Application & Calculation
Measure Engine: Assigns Numerator/Denominator based on LLM-classified data.

Key Idea:

LLM Agents use structured and unstructured data (e.g., pulling 'PHQ-9 = 12' from narrative to infer 'positive screen,' or identifying a documented counseling session as follow-up) to classify patients according to CQM criteria. These classifications then feed into a measure engine. The measures themselves don't need to be written in a 'computable' way that assumes specific codes/fields are populated. This makes measurement resilient to variations in EHR documentation practices.

Unlocking True Value: Advantages of the New Approach

This flexible, post-EHR computation architecture directly supports a scalable, learning health system infrastructure.

✅ Works across all EHRs (no per-vendor custom build).
✅ Zero extra clicks for clinicians at point of care.
✅ New or "what-if" measures calculable retrospectively.
✅ Same robust pipeline for "real" care analysis & "report-only" pilots.
✅ More accurate reflection of true clinical care.
✅ Reduces clinician burnout related to quality reporting.

Key Policy Levers

These steps can shift the cost and burden of quality measurement away from the point of care and toward shared, efficient services.

  1. Require Bulk Export performance to match vendor-proprietary methods.
  2. Require an API to automate Full EHI Export (Electronic Health Information).
  3. Ensure health systems are contractually allowed to drive EHR interactions with automated agents (address info blocking).
  4. Deploy a richer set of "report only" measures to push on system flexibility!