Discover the pros and cons of AI-augmented risk adjustment and how tech + expertise drive results.

Improving Accuracy & Data Integrity

Defensible, Audit-Ready Records

Automating Clinical Documentation

Precise Coding Across Care Settings

Complete Coding for Ancillary Services

Optimized Codes for Proper Reimbursement

Protecting Revenue Through Coding

Optimizing RAF for Population Health

Analytics-Driven Risk Adjustment

Improving Risk Capture Accuracy

Real-Time Coding for Better Outcomes

Accurate Data From First Touch

Preventing Delays Before Care

Recovering Revenue From Denials

Accelerating Payer Responses

Capturing Charges Without Leakage

Reducing Claim Errors Early

Resolving Credits With Precision

Accurate Payments, Faster Close

Strengthening Payer Appeals

Improving Accuracy Through Expert Audits

Compliance & Risk-Based Training

Risk-Focused Documentation Compliance

Compliance & Risk-Based Training

Risk-Focused Documentation Compliance

Importance of Human Oversight in AI-Assisted Medical Coding Backed by Real-World Evidence

Human oversight in AI-assisted medical coding is not redundant. It is the mechanism that catches what automation misses. Based on real-world data gathered from 40,211 outpatient encounters at two facilities of a large U.S. health system, Chirok Health found that their certified clinical coders with clinical backgrounds corrected approximately three in ten charts even after an AI-assisted coding tool and provider had already reached a determination. That figure, replicated independently across two separate facilities in the same month, is the real-world answer to how much human review adds after AI-assisted coding is already in place.

The corrections were not a formality. Chirok Health’s coding experts raised E/M levels on charts where the AI coding tool under-leveled the encounter, lowered them where the documentation did not support the assigned code, and reclassified service types on roughly 8 percent of encounters where the tool applied E/M coding to a visit that required a different code set entirely. Both samples produced the same correction rate. The data does not make a theoretical argument for human oversight in AI-assisted medical coding. It measures it directly, at scale, under real operating conditions.

AI medical coding audit

Why AI-Assisted Coding is Used in Healthcare Organizations?

The Operational Drivers Behind Automation Adoption

The volume of outpatient encounters in most health systems makes manual first-pass E/M leveling difficult to sustain at scale. A high-volume outpatient operation documenting hundreds of encounters daily generates a coding workload that staffing alone cannot reliably absorb without delays, variability, or error. AI-assisted medical coding tools were developed to address that throughput problem. They apply documentation guidelines logic to the elements of an encounter and return a suggested level quickly, reducing inconsistency across large encounter volumes.

How AI-Assisted Medical Coding Tools Function in Practice

Most AI-assisted coding systems evaluate structured documentation elements: history components, examination findings, and the complexity of medical decision-making as recorded by the provider. Some platforms also parse free-text documentation using natural language processing to identify clinical elements that structured fields do not capture. The system returns a suggested E/M level; the provider reviews, adjusts, or accepts it. That combined output, the AI coding suggestion filtered through provider selection, is the baseline Chirok Health measured human review against. The tool and the provider had already reached a determination before the Chirok coder looked at the chart.

Why Human Oversight Remains a Non-Negotiable Layer

The Limits of AI-Assisted Coding Without Clinical Judgment

AI-assisted medical coding tools are calibrated to the documentation the provider enters. They apply guidelines logic to the inputs they receive; they do not evaluate whether those inputs are clinically consistent with the encounter narrative, whether the service type recorded matches the documented visit, or whether payer-specific requirements introduce coding conventions that the AI coding logic does not account for. A preventive visit coded as a standard office visit is not an acuity error. It is a type error, and the AI tool has no basis to flag it because it is operating within the framework the provider established.

Compliance Exposure When Human Review Is Removed

The Office of Inspector General (OIG) has identified improper E/M coding as a recurring source of Medicare overpayments through its annual work plan reviews and the Comprehensive Error Rate Testing (CERT) program. When a facility relies exclusively on automated leveling plus provider confirmation, errors that a trained coder would catch at the pre-bill stage go undetected. By the time a payer audit or retrospective review surfaces them, the claim window has typically closed, and recoupment risk replaces what would have been a straightforward pre-submission correction.

What Secondary Review Actually Adds After Automation

Secondary review by a certified clinical coder is not a second pass of what the automated tool already did. The coder brings clinical judgment; the tool is not designed to replicate: evaluation of whether the documentation supports the assigned level under AMA guidelines, awareness of payer-specific coding conventions, and the ability to identify service type mismatches that fall outside the E/M framework entirely. The Chirok data measures this precisely. The tool and the provider had already applied their combined logic. The coder’s independent evaluation then identified where that logic produced an incorrect result.

What the Data Actually Shows

How the Study Was Designed

Chirok Health reviewed 40,211 outpatient encounters from two facilities of the same U.S. health system over a single calendar month. Each encounter had been leveled by an automated E/M calculator, and a provider selection had been recorded. Chirok clinical coders then reviewed each chart independently, comparing their determination against the tool-plus-provider output. Reviewers were not blinded to the automated level before reviewing the documentation, which introduces the possibility of anchoring bias. That limitation works against the finding rather than for it: a reviewer who sees the original determination is more likely to leave it unchanged. The correction rates reported here are a floor, not a ceiling.

Facility 1 - Results of 18,293 Encounters

Coder Action Encounters % of Reviewed
E/M level left unchanged 12,382 67.7%
E/M level raised by Chirok clinical coder 2,555 14.0%
E/M level lowered by Chirok clinical coder 1,843 10.1%
Service type corrected by Chirok clinical coder 1,513 8.3%
Total corrected 5,911 32.3%

Facility 2 - Results of 21,918 Encounters

Coder Action Encounters % of Reviewed
E/M level left unchanged 15,379 70.2%
E/M level raised by Chirok clinical coder 2,974 13.6%
E/M level lowered by Chirok clinical coder 1,862 8.5%
Service type corrected by Chirok clinical coder 1,703 7.8%
Total corrected 6,539 29.8%

Breaking Down What Coders Are Correcting and Why It Matters

E/M Level Adjustments Run in Both Directions

On encounters where E/M leveling applied, coders corrected the tool-assisted level on 26 percent of charts in Facility 1 and 24 percent in Facility 2. Those corrections were not uniformly upward. Coders raised levels more often than they lowered them (14.0 vs. 10.1 percent in Facility 1; 13.6 vs. 8.5 percent in Facility 2), but the lowering rate is not trivial. A facility operating without secondary review is not only failing to capture revenue at the correct level; it is also submitting codes that in some cases exceed what the documentation supports. That is an over-coding exposure, and it carries a different set of compliance consequences than under-coding.

Service Type Corrections Are a Separate Risk Category

The 8.3 percent service type correction rate in Facility 1 and 7.8 percent in Facility 2 describe a distinct error category. These are not acuity miscalculations where a Level 3 should have been a Level 4. These are encounters where E/M leveling did not apply because the documented service was a preventive exam, a consultation, an admission, or another visit type governed by a different coding pathway. When an automated tool assigns an office visit E/M code to a preventive service encounter, the downstream effects extend beyond revenue: payer adjudication rules differ by service type, benefit category implications differ, and the compliance exposure under a post-payment audit differs as well.

Why the Correction Rates Are More Likely a Floor Than a Ceiling

The study design carried a known methodological limit: Chirok clinical coders were not blinded to the automated level before reviewing the chart documentation. Research on anchoring effects in clinical and judgment settings consistently finds that reviewers who see a prior determination are statistically more likely to preserve it rather than deviate from it. That means the 32.3 and 29.8 percent correction rates here reflect a conservative count. Under blinded conditions, where coders evaluate documentation without first seeing the automated output, correction rates would plausibly be higher.

The Compliance and Revenue Integrity Implications

The Cost of Uncorrected Errors

The financial consequence of an uncorrected E/M coding error depends on its direction. An upward correction represents revenue the facility earned but did not fully capture; the documentation supported a higher level than the tool assigned. A downward correction means the submitted code exceeded what the documentation supports, which is an overpayment risk. The OIG’s CERT program has consistently identified E/M services as a category with elevated improper payment rates in Medicare fee-for-service. When coding corrections run at roughly 30 percent across a high-volume outpatient operation, the aggregate financial and compliance exposure is material regardless of the direction any single error takes.

E M coding analysis

What This Means for Payer Adjudication and Audit Risk

Payers apply edit logic and medical necessity criteria at the service-type level, not just the acuity level. A claim coded as an office visit when the documented service was preventive may clear an automated clearinghouse check and still fail a post-payment audit because the benefit category is wrong. Recovery Audit Contractors (RACs) have targeted outpatient E/M coding patterns consistently since the program expanded to include outpatient claims. Organizations without a pre-bill secondary review process are more exposed when those audits arrive, because they have no mechanism to catch service-type errors before the claim reaches the payer.

Why Automated Tools and Human Review Are Not Interchangeable

What AI-Assisted Coding Does Well

AI-assisted medical coding tools solve a real operational problem. They apply consistent documentation guidelines logic to large encounter volumes quickly, reduce variability in how providers self-assign levels, and create an auditable record of the leveling rationale. The Chirok Health’s study was designed specifically to measure the value human review adds after the AI coding tool has already performed its function. The automated output was the baseline, not the subject being tested. Nothing in these findings argues against using AI-assisted coding tools; they argue for adding a secondary review layer on top of them.

Where Human Clinical Judgment Is Still Required

The gap between what an AI-assisted coding tool does and what a certified clinical coder does is not primarily a matter of processing speed or data volume. The coder brings clinical context, payer-specific knowledge, and the ability to assess whether a code assignment is defensible under audit conditions. When documentation is technically complete but the service type recorded does not match the encounter narrative, the AI coding tool has no basis to flag the discrepancy. When payer-specific coding conventions deviate from AMA guidelines, the tool has no mechanism to apply them. Both error categories appeared in the Chirok Health’s data. The AI-assisted coding tool handled what it was designed to handle. The coder caught what it was not.

Do you want a copy of this data and analysis?

If you are using AI for coding and need to discuss this internally, we have created a one-pager that contains the data in an infographic.

Download the Analysis

What 40,211 Encounters Confirm About Human Oversight in Coding

The argument for human oversight in AI-assisted medical coding does not rest on a theoretical limitation of automated tools. It rests on what happened when Chirok Health placed a certified clinical coder in the review seat after the tool and the provider had already made their determination: three in ten charts required a correction, the pattern held across two independent samples from the same health system, and the corrections ran in both directions. A better tool would not close that gap, because the gap is not a tool problem. It is a clinical judgment problem. And the data, 40,211 encounters across two facilities in a single month, says the same thing each time.

FAQs

What Is the Difference Between Automated E/M Leveling and Clinical Coding Review?

Automated E/M leveling applies documentation guideline logic to the structured elements of an encounter and suggests an E/M code level. Clinical coding review is performed by a certified coder who verifies whether the documentation supports the assigned level, confirms the correct service type, and applies payer-specific coding rules. Automated tools process documentation, while clinical reviewers apply professional judgment. Chirok Health’s study measured the gap between these two functions.

How Often Do Automated E/M Tools Produce an Incorrect Level or Service Type?

In the Chirok Health study of 40,211 encounters, the combined correction rate after automated E/M leveling and provider confirmation was 32.3 percent in the first sample and 29.8 percent in the second. On encounters where E/M leveling applied, the correction rate was 26 percent and 24 percent respectively. Those figures include both upward and downward adjustments. The service type correction rate, where the encounter required a different code set rather than a different acuity level, was approximately 8 percent in each sample.

What Compliance Risk Does an Incorrect E/M Code Create?

An E/M code above what the documentation supports is an overpayment risk under Medicare and most commercial payer contracts. The OIG's CERT program and Recovery Audit Contractor reviews have identified E/M coding as a high-error category in Medicare fee-for-service. An E/M code below what documentation supports is a revenue loss. A service type error, where a preventive visit is coded as an office visit or a consultation is entered as a follow-up, affects benefit adjudication, patient cost-sharing calculation, and payer audit outcomes, independent of the acuity level assigned.

Does Adding Secondary Coding Review Slow Down Claim Submission?

A well-structured concurrent coding review program does not require holding claims beyond the normal billing cycle. Chirok Health's review model operates within the standard pre-bill window, identifying corrections before submission rather than through retrospective audit. The correction window before submission is where the financial and compliance value is concentrated. Post-submission corrections require claim adjustments or appeals, which carry their own time and administrative cost that pre-bill review avoids.

What Service Types Are Most Commonly Misclassified by Automated E/M Tools?

Preventive visits, consultations, and admission-day services are the categories where service type misclassification most frequently occurs in high-volume outpatient settings. These encounter types follow coding pathways that differ from standard office visit E/M guidelines, and automated tools calibrated to the standard E/M framework may not flag when a documented service belongs to a different category. The study did not break out service type corrections by category, but the consistent 8 percent rate across both samples reflects a pattern across a diverse outpatient encounter mix.

When Should a Health System Add a Secondary Clinical Coding Review Program?

Secondary review adds measurable value at any encounter volume large enough that coding errors accumulate into material financial and compliance exposure. For health systems operating above 5,000 outpatient encounters per month, a 25 to 30 percent correction rate on automated E/M outputs represents a volume of errors that a structured pre-bill review process can systematically address. The decision is not whether errors are occurring at that scale; the Chirok data confirms they are. The decision is whether the current workflow has a mechanism to catch them before they reach the payer.

Author Bio:

Kanar Kokoy

CEO - Chirok Health

Healthcare CEO & CDI/RCM innovator. I help orgs boost accuracy, integrity & revenue via truthful clinical docs. I've led transformations in CDI, coding, AI solutions, audits & VBC for health systems, ACOs & more. Let's connect to modernize workflows.

Table of Contents