
The AI Billing Arms Race: Who Will Win?
Our patients may look more “complex” on the claim but are they actually sicker?
A new Blue Cross Blue Shield Association analysis found that hospitals are coding more inpatient stays as complex, without a corresponding increase in selected measures of treatment intensity. BCBS points to AI-assisted coding as a possible driver, but the data can’t tell us whether every additional diagnosis reflects better documentation or overstated complexity.
I don’t think we have a clear answer yet. More complete documentation can help us physicians get paid appropriately for care we already provide. But it doesn’t, by itself, show that we’re delivering better care at a lower cost.
The same distinction applies to insurers. AI can help them scrutinize claims and challenge payment faster without necessarily making those decisions more clinically sound.
This is the AI billing arms race: hospitals invest in tools to capture and defend revenue, while insurers invest in tools to scrutinize and contest those claims. Both sides can improve their bottom line without demonstrating a benefit to patients.
In this Huddle Trends report, I’ll break down the new BCBS data, how both sides are building their AI capabilities, and what would need to change for this race to produce more than better billing and faster disputes.
The New Data: More Complexity, Similar Care
I covered an earlier BCBS analysis in How AI Documentation Tools Are Making Upcoding Worse. That report highlighted rising postpartum anemia coding without a corresponding rise in transfusions. This new September white paper examines the bigger inpatient trend, then digs into major bowel procedures.
First, a quick billing explanation (which you can learn more about in my Healthcare 101 course).
Hospitals generally receive a bundled payment for an inpatient stay based on its diagnosis-related group, or DRG.
Certain secondary diagnoses can move the same admission into a higher-paying category because they signal greater complexity and expected resource needs. For example:
DRG 193: Simple pneumonia and pleurisy with MCC (Major Complication or Comorbidity).
DRG 194: Simple pneumonia and pleurisy with CC (Complication or Comorbidity).
So documenting acute blood-loss anemia alongside the simple pneumonia and pleurisy can change reimbursement, even if the management itself is unchanged. BCBS calls these diagnoses “bump codes.”
Across the Blue system’s inpatient claims, the share classified as complex rose from roughly 37% in early 2023 to 40% by late 2025. BCBS estimates that increased coding intensity added $942 million in claims costs over two years. About $653 million—roughly 70%—was associated with secondary diagnoses moving cases into higher-paying categories (that’s the “bump coding”), representing 55,158 additional complex cases relative to its 2023 baseline.
These are BCBS’s estimates of incremental payments. They are not an audited tally of fraudulent claims or a direct measurement of spending caused by AI. But, this is all estimated in the context of AI coding tools becoming more prominent.
The Bowel Surgery Example
For major bowel procedures, the share of claims in the highest-complexity category increased from 20.2% to 22.7% between 2023 and 2025. BCBS attributes $60.8 million in incremental claims costs to coding shifts in this DRG family.
The diagnoses driving those shifts included acidosis, hyponatremia, acute posthemorrhagic anemia, and malnutrition. These are familiar conditions to us physicians—and conditions that software scanning labs and notes might help identify. An abnormal lab value alone, though, does not establish a clinically meaningful, reportable diagnosis (a patient who’s sodium is 133 from a baseline of 136 is technically “hyponatremic” but you wouldn’t necessarily act on it because the sodium is more or less at baseline).

Source: BCBS
The more interesting comparison is between hospitals. In 2025, hospitals in the top quarter of coding-complexity growth classified 75.6% of bowel cases as complex, compared with 65.0% at other hospitals. Yet several measures of resource use were similar or lower:
ICU use: 11.5% versus 13.2%.
Transfusions: 3.6% versus 3.9%.
Median length of stay: four days in both groups.
The overall mismatch is what BCBS calls “clinical discordance”: more complexity on the claim without a comparable increase in selected measures of care. Anemia, again, is a good example. High-growth hospitals coded posthemorrhagic anemia in 13.7% of cases versus 9.9% elsewhere, but among patients with that diagnosis, transfusion rates were lower: 16.9% versus 19.3%.
None of this was analyzed for significant differences, there is no causality—just a correlation. So be a bit skeptical, especially since BCBS has a financial stake in this story. Insurers have a reason to publicize rising hospital payments, just as hospitals have a reason to publicize denials and delayed reimbursement. Additionally, more complete coding could also be correcting longstanding underdocumentation. If we were treating these conditions all along but failing to capture them, payment could increase without treatment changing. That is different from adding unsupported diagnoses to inflate a bill.
So, BCBS identifies a “relationship” worth investigating further, but it does not establish how much reflects better documentation, true different patients, or inappropriate coding—or how much AI caused. Sorting that out requires clinical record review and a stronger comparison of adopters and non-adopters.
The Conundrum: Both Sides Have a Reason to Invest
From a hospital’s perspective, investing in AI makes sense to ensure the care provided is translated into documentation, codes, and claims that meet the payer’s rules for reimbursement….





