Discover the pros and cons of AI-augmented risk adjustment and how tech + expertise drive results.

Improving Accuracy & Data Integrity

Defensible, Audit-Ready Records

Automating Clinical Documentation

Precise Coding Across Care Settings

Complete Coding for Ancillary Services

Optimized Codes for Proper Reimbursement

Protecting Revenue Through Coding

Optimizing RAF for Population Health

Analytics-Driven Risk Adjustment

Improving Risk Capture Accuracy

Real-Time Coding for Better Outcomes

Accurate Data From First Touch

Preventing Delays Before Care

Recovering Revenue From Denials

Accelerating Payer Responses

Capturing Charges Without Leakage

Reducing Claim Errors Early

Resolving Credits With Precision

Accurate Payments, Faster Close

Strengthening Payer Appeals

Improving Accuracy Through Expert Audits

Compliance & Risk-Based Training

Risk-Focused Documentation Compliance

Compliance & Risk-Based Training

Risk-Focused Documentation Compliance

Human Oversight in AI-Assisted Medical Coding

Purpose

As organizations evaluate AI-assisted and autonomous coding, the question before leadership is not whether to adopt AI, but how to govern it. This briefing sets out why a human-in-the-loop standard should be a condition of any coding-automation deployment rather than an optional refinement. It draws on three things: the current independent research on AI coding accuracy, the emerging regulatory record, and an operational analysis Chirok Health conducted on more than 40,000 encounters at a health system already using an automated AI (Evaluation and Management) E/M coding tool. 

The opportunity is real

AI-assisted coding can accelerate throughput and reduce routine error. Manual coding is itself imperfect, with published inaccuracy commonly in the range of 10 to 20 percent of cases, and well-designed automation can help close that gap on clean, unambiguous charts. 

What the evidence shows about accuracy

Independent research finds that general-purpose AI remains unreliable at coding on its own. In a Mount Sinai benchmark published in NEJM AI, the best-performing model, GPT-4, produced exact-match ICD-10-CM codes only 33.9 percent of the time, with every model evaluated scoring below 50 percent. The authors concluded these tools are not suitable for direct use in medical coding without further refinement. Accuracy improves substantially with retrieval techniques and domain-specific fine-tuning, yet even purpose-built models reach roughly 69 percent exact match on real-world clinical notes, and error rates remain non-trivial on the complex encounters and low-prevalence codes that are hardest to assign. The pattern across the literature is consistent: AI is strongest on routine, bounded cases and weakest precisely on the complex, high-acuity documentation that is most consequential to code correctly. 

In bounded settings such as the emergency department, retrieval-augmented models reviewed by clinicians have matched or exceeded provider-assigned codes for accuracy and specificity. 

The professional bodies that set standards for this work reach the same conclusion from practice rather than the laboratory. American Health Information Management Association, AHIMA, the coding profession’s standards organization, reports that routine audits sampling only one to two percent of cases may show roughly 95 percent accuracy, while a targeted audit of higher-risk work can reveal errors in 30 to 40 percent of the cases reviewed,5 and it holds that strong human oversight, especially through audits, remains essential as AI enters the workflow. The Association of Clinical Docu[mentation Integrity Specialists, ACDIS, is more pointed: even advanced systems can hallucinate or overreach, for instance surfacing a long-resolved condition as active or inferring a diagnosis the record does not support, so human verification is, in its words, a “professional guardrail, not a suggestion.”

AI medical coding audit

The stakes: revenue and compliance

The charts AI handles least well are the ones that carry the most value and the most risk. Complex, high-acuity encounters concentrate both the work-RVU and risk-adjustment value of a system’s revenue and a large share of its audit exposure. Coding such charts only to the level sufficient to submit a clean claim under-captures earned revenue; coding them upward without clinician review creates the opposite exposure. The coding profession’s own guidance warns that AI can inadvertently drive both upcoding and downcoding when proper checks are absent.5 Provider coding already operates inside a mature audit regime, Risk Adjustment Data Validation in Medicare Advantage and False Claims Act liability in fee-for-service, and unreviewed automated coding at scale concentrates exactly the risk those regimes exist to detect. The financial exposure runs in both directions, and human review is the single control that addresses both.

Real-World Validation: Measuring Incremental Value of Human Review

Chirok Health analyzed 40,211 encounters from two facilities of a large U.S. health system in a single month, each leveled with the automated E/M calculator and a provider selection recorded, then reviewed by a Chirok clinical coder. The comparison is the tool-plus-provider selection against the coder’s independent determination, so it measures the value human review adds after an automated level is already in place, rather than AI in isolation. The pattern was held independently in both facilities.  

Facility 1 (18,293 encounters reviewed)

Coder Action Encounters % of Reviewed
E/M level left unchanged 12,382 67.7%
E/M level raised by Chirok clinical coder 2,555 14.0%
E/M level lowered by Chirok clinical coder 1,843 10.1%
Service type corrected by Chirok clinical coder 1,513 8.3%
Total corrected 5,911 32.3%

Facility 2 (21,918 encounters reviewed)

Coder Action Encounters % of Reviewed
E/M level left unchanged 15,379 70.2%
E/M level raised by Chirok clinical coder 2,974 13.6%
E/M level lowered by Chirok clinical coder 1,862 8.5%
Service type corrected by Chirok clinical coder 1,703 7.8%
Total corrected 6,539 29.8%

On the encounters where E/M leveling applied, coders corrected the tool-assisted level on roughly one in four (26 percent at the first sample, 24 percent at the second), and the corrections ran in both directions rather than uniformly raising acuity. On a further 8 percent of charts, E/M leveling did not apply, because the encounter was a different kind of service, such as a preventive visit, consult, admission, etc. and coders corrected the code to match the documented services. This is an accuracy fix to the type of service and can result in an acuity, revenue and/or compliance change. In total, the review corrected about three in ten charts. Each level was assessed against the encounter documentation; reviewers were not blinded to the original level, so some anchoring to it cannot be excluded. If anything, that works against the finding rather than for it: a reviewer who sees the original level tends to leave it in place, so the correction rates here are more plausibly a floor than a ceiling. 

Because human coders corrected a meaningful share of charts even with an advanced automated E/M tool in place, these findings provide real-world evidence that secondary clinical coding review is a quality, compliance, and revenue-integrity safeguard rather than a redundant step in the coding process.

Human oversight is becoming the standard

Regulators have begun writing human oversight of healthcare AI into law, and the first target is the payer side. Effective July 2026, HB 481 bars a health insurance carrier from making an adverse determination on a prior-authorization request, for prescription drugs or health care services, unless a licensed physician has reviewed and approved it. The requirement applies whether AI was involved or not, and it reaches fully-insured plans, with Medicare, Medicaid, and self-funded ERISA plans outside its scope. AI-specific measures are moving alongside it: in the 2026 session, SB 586, which would require carriers to disclose their AI use to state regulators, bar determinations made solely by AI, and provide expedited external review, passed the Senate but was continued by the House to the 2027 session, and a state legislative technology commission endorsed a set of healthcare-AI recommendations in late 2025, including provider transparency and internal-standards requirements.

This state activity is now meeting federal resistance: a December 2025 executive order directs agencies to challenge state AI laws, and CMS has begun testing AI-assisted prior authorization in traditional Medicare. The near-term regulatory picture is contested, not settled, but the case for human review does not depend on how that contest resolves: the standard being written on the payer side is portable, and the bodies that accredit providers are moving the same way.  

In 2025 the Joint Commission, with the Coalition for Health AI, issued guidance on the responsible use of AI in healthcare that sets governance and oversight expectations, including transparency. For a provider, the more immediate pressure is not a pending statute at all. Inaccurate automated coding already creates False Claims Act and RADV exposure, and that exposure turns on whether the organization can demonstrate human review of its coding. An organization that has automated its coding but cannot show that review is not waiting on a future law; it is already carrying the risk that both auditors and regulators increasingly expect it to control.

Recommendation

Adopt AI coding for scale and make human-in-the-loop review a condition of deployment rather than an afterthought. Leadership should require that any platform your organization selects: 

  • routes complex and high-value encounters to qualified human review rather than auto-finalizing them; 
  • supports concurrent documentation review before the claim is released; and 
  • produces an auditable record of human oversight for consequential coding decisions. 

These requirements are platform-agnostic. They preserve the efficiency AI delivers on routine volume while protecting the revenue and the compliance posture that depend on getting the complex cases right. 

Sound governance also means setting the standard before deployment, not after. That means measuring current human coding accuracy on the complex, high-value charts before AI is introduced, validating the system in parallel before it finalizes any claim, auditing on a risk-weighted basis rather than by routine low-percentage sampling, and keeping human review active rather than a sign-off on the machine’s suggestion. These steps let the organization prove, rather than assume, that automation is improving accuracy and revenue integrity rather than quietly eroding them.

E M coding analysis

Author Bio:

Brandy Kerby

CGO - Chirok Health

Brandy Kerby is Chief Growth Officer at Chirok Health, a clinician-led medical coding, CDI, and revenue cycle organization. She works with health systems, providers, plans, and risk-bearing organization leaders on a question most are now facing: not whether to adopt coding automation, but how to govern it. Her work centers on the operational evidence behind that decision, including the encounter-level analysis presented in this briefing.

Table of Contents