top of page
  • LinkedIn
  • Instagram
  • Facebook
  • X

What 238 Business School Cases Reveal About AI's Readiness for Knowledge Work

AI in business school
What 238 Business School Cases Reveal About AI's Readiness for Knowledge Work

Frontier AI models already score above 87 percent against instructor-written grading standards on hundreds of open-ended business school case questions, according to a new working paper from researchers at Wharton, Harvard Business School, Carnegie Mellon's Heinz College, and the University of Pennsylvania's Department of Computer and Information Science. The paper, titled "Frontier AI performance across the business disciplines," introduces BusinessCaseBench, a benchmark built from 238 licensed case studies spanning eighteen business disciplines, and it arrives at a moment when most AI benchmarks still measure narrow, verifiable tasks like factual recall, math problems, or code that either compiles or doesn't.


A business school case does not fit that narrow, verifiable-answer approach. It hands a student a messy narrative about a real or fictional company, drops in ambiguous evidence and irrelevant details, and asks for a defensible recommendation under incomplete information. There is often no single correct answer, only a range of analyses an instructor would accept. That is precisely the kind of work the paper's authors, Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, and Karim Lakhani, argue current AI benchmarks fail to capture, and it also fills the daily routine of financial analysts, consultants, and mid-level managers.


Borrowing a Method Medicine Already Validated


The researchers point to a precedent from clinical AI research. High scores on the multiple-choice United States Medical Licensing Examination led many observers to assume language models had matched real clinical reasoning, but a 2025 study instead tested models on New England Journal of Medicine case challenges, open-ended diagnostic puzzles written for practicing physicians. Models performed strongly there too, and the result carried more weight because the format mirrored real practice rather than an administrative exam.


BusinessCaseBench applies the same logic to management education. Business school cases are curated, expert-graded, and usually gated behind publisher licenses and instructor-only solution guides, which limits how freely the benchmark can be shared but also lowers the odds that models memorized the answers during training. The researchers ran a contamination check across several major web-crawled training datasets, including C4, The Pile, RedPajama, Dolma, and DCLM, and found no evidence that the case text or solutions appear verbatim in any of them.


Turning Case Solutions Into Grading Rubrics


Each of the 615 benchmark questions was built from a real instructor case solution rather than inferred from the case narrative itself, then converted into an equally weighted checklist of criteria so a model answer earns credit item by item rather than receiving one holistic grade. A model reads the full case and question in a single prompt, with no outside tools or follow-up clarification, and a separate AI system acting as judge scores the answer against the checklist and the reference solution. Three trained human annotators independently built their own rubrics and grades on a sample of questions to check whether the automated process tracked expert judgment. The correlation between human and automated scores was moderate but directionally consistent, and annotators rated the great majority of automated rubrics and grades as acceptable.


The benchmark spans questions built on real firms and fictional ones, numerical and non-numerical prompts, and subjective and objective framing, and every question is also mapped to the Department of Labor's O*NET occupational classification system. That last step lets the results be read against actual job categories rather than just academic subject lines.


High Scores Under Partial Credit


Across all 615 questions, three current frontier models, Anthropic's Claude Sonnet 4.6, OpenAI's GPT-5.4, and Google's Gemini 3 Flash Preview, each performed at a consistently high level. Claude Sonnet 4.6 led at 88.4 percent, GPT-5.4 followed at 87.2 percent, and Gemini 3 Flash Preview came in at 81.6 percent. The gap between the top and bottom of that range was only 6.8 percentage points, and confidence intervals for the two leading models overlapped throughout, meaning the paper stops short of declaring a clear winner among them.


Under this scoring approach, called Standard scoring in the paper, a model earns credit for every rubric item it satisfies, even if the overall answer misses other pieces of the expected analysis. That is a reasonable way to measure how much of an instructor's expected solution shows up in a response, and by that measure, frontier models already cover most of what a case instructor would look for.


A Stricter Bar Tells a Different Story


The paper also reports a second, stricter metric called Complete Answer scoring, which only counts a response as successful if it satisfies every single rubric criterion on a question. Under that bar, Claude Sonnet 4.6 dropped to 49.6 percent, GPT-5.4 to 47.6 percent, and Gemini 3 Flash Preview to 32.0 percent. Even the strongest model left more than half its answers incomplete by instructor standards, and the spread between models widened to 17.6 percentage points, roughly two and a half times the gap seen under partial credit.


That divergence is the paper's central finding. A model can produce an analytically strong draft that would earn a solid grade under partial credit while still missing a required element, whether that is a specific number, a named risk, or a clear final recommendation. The authors describe this pattern as AI output functioning more like a draft awaiting review than a finished verdict, and they treat the stricter metric as a deliberately conservative lower bound rather than the final word on model quality, since many instructors would treat a high partial-credit answer as analytically sufficient on genuinely open-ended questions.


Where the Eighteen Disciplines Diverge


Discipline mattered more than any other factor the researchers tested. Standard scores ranged from 80.1 percent in Marketing and Sales up to 95.0 percent in Business and Government Relations, a wider spread than the gap between the three models themselves. Under the stricter Complete Answer scoring, the same disciplines separated even further, from 25.8 percent in Operations and Service Management to 82.5 percent in Decision Analysis. All three models ranked the disciplines in nearly the same order, which suggests the difficulty comes from the structure of the task itself rather than from any one company's training approach.


A regression analysis in the paper found that discipline and question type together explained only about 5 percent of the variation in how well models scored on individual questions. Knowing which specific case a question came from explained far more, about 22 percent, which suggests difficulty is largely a property of the individual case rather than a broad category a business school or hiring manager could point to in advance.


Real Firms, Fictional Firms, and a Finance Outlier


Questions built on fictional companies scored slightly higher on average than questions built on real ones, by 1.9 percentage points, but that pattern masks a sharp reversal in Finance, where fictional-company questions scored 15 percentage points lower than questions about real companies, with Accounting showing a smaller version of the same gap. The researchers suspect real, named companies sometimes have financial details that already circulate in public filings or financial press coverage, giving models a head start that invented companies cannot provide.


What Defeats These Models, and What Doesn't


Very few questions stumped every model tested. Only 43 of the 615 questions, about 7 percent, produced a top score of 70 percent or below across all three frontier models, and only about 6 percent of individual rubric criteria went unmet by every model. When the researchers built a hypothetical "oracle" that took the best of the three models' scores on each question, aggregate performance rose to 92.8 percent, 4.5 points above the strongest single model. What one model misses on a given question, another model frequently catches, which points toward incomplete coverage of multi-part rubrics as the main source of remaining error rather than a fundamental gap in business knowledge.


Mapped onto O*NET work activities, the pattern held up at the level of specific job tasks. Structured, well-defined activities like explaining financial information or calculating financial data approached ceiling-level scores. Open-ended advisory work, such as identifying business or organizational opportunities or advising others on financial matters, ranked among the hardest activities for every model tested.


Two Years of Rapid Movement Within One Model Family


To trace how quickly this capability has developed, the researchers evaluated four successive OpenAI releases on the identical question set, spanning roughly two years from GPT-4 Turbo through GPT-4.1, GPT-5, and GPT-5.4. Standard scores climbed from 63.9 percent to 87.2 percent, a 23.3 percentage point gain, while the stricter Complete Answer score climbed further in relative terms, from 13.2 percent to 47.6 percent. The gains showed up broadly rather than concentrating in the numerical, quantitative tasks early language models struggled with most and where popular benchmarks tend to focus. Non-numerical and subjective questions improved by as much or more than numerical and objective ones.


The paper also reports an after-the-fact extension covering Claude Fable 5, an Anthropic model released after the primary experiments were complete, whose rollout was briefly interrupted by U.S. export controls tied to cybersecurity concerns before access was restored. Its performance came in close to Claude Sonnet 4.6 on both metrics, which the authors attribute to substantial overlap in training data between successive models from the same provider rather than to any regression in capability.


What This Means for Business Schools and Early-Career Work


Case pedagogy exists to train synthesis, judgment under uncertainty, and defensible reasoning, the skills that have historically anchored early-career analytical roles in consulting, finance, and general management. With frontier models already producing strong drafts on this work and improving quickly within at least one model family, the cost of generating a plausible first-pass analysis is falling fast.


The design challenge for business education, in the authors' reading, shifts from teaching students to produce a competent analysis toward teaching them to verify, complete, and recognize what a fully adequate answer requires, especially in disciplines where partial credit already runs high but complete answers remain rare. For labor markets, the O*NET mapping offers a rough guide to where entry-level analytical tasks look most exposed and where human judgment may retain an edge, particularly in work built around integrating stakeholders' interests, exercising accountability, or navigating live organizational politics, capabilities that fall outside a single-turn benchmark like this one.


Where the Study Draws Its Own Limits


The authors are direct about what this benchmark does not capture. Every evaluated interaction is single-turn, with no back-and-forth clarification, tool use, or negotiation, so the results say nothing about how these models handle the iterative, multi-turn work real jobs often involve. The benchmark is English-only, and the automated grading system, while checked against human annotators, still carries the general risks of using one AI model to judge another's output. The researchers also note that while they found no sign of the case materials in major public training datasets, frontier models are trained on proprietary data whose full contents cannot be independently verified.


Those caveats point toward the paper's own suggested next step: pairing this kind of case-grounded benchmark with studies of how AI performs in live, multi-turn, organizational settings, where the gap between a strong draft and a fully accountable decision is likely to matter even more than it does here.

Author Bio


David Borish is the author of the forthcoming book The Tony Hawk Paradox, which examines how AI capabilities first proven in controlled or simulated environments consistently transfer into broader real-world systems. He writes long-form analysis on frontier AI research, enterprise deployment, and technology policy at davidborish.com.


The Tony Hawk Paradox Book
Click image to learn more

 
 

JOIN THE AI SPECTATOR MAILING LIST

CONTACT

Contacting You About:

Thanks for submitting!

New York, NY           

Db @DavidBorish.com           

  • LinkedIn
  • Instagram
  • Facebook
  • X
Back to top

© 2026 by David Borish IP, LLC, All Rights Reserved

bottom of page