top of page

Latham & Watkins Goes Off the Cloud: Big Law's First Embrace of the Open-Prem Framework

7 minutes ago
6 min read
Latham & Watkins Goes Off the Cloud: Big Law's First Embrace of the Open-Prem Framework
Latham & Watkins Goes Off the Cloud: Big Law's First Embrace of the Open-Prem Framework

What Latham Actually Bought


Latham & Watkins, the United States' second-largest law firm by revenue at $8.3 billion last year, has purchased several servers loaded with multiple Nvidia GPUs and is using them to run and customize open-weight AI models inside its own systems, according to Financial Times reporting picked up this week by Law360 and Legal IT Insider. The firm is described as the first major law firm publicly known to build this kind of infrastructure itself, rather than subscribing to legal AI platforms like Harvey or Legora, or building applications on top of models it still rents from a hyperscaler. That distinction matters: most large firms adopting generative AI have layered their own workflows on top of someone else's compute and someone else's model weights, while Latham now owns the GPUs and fine-tunes open-weight models on them directly, supported by an internal technology organization large enough to do the work in-house rather than through a vendor.


Rene Mendoza, Latham's chief information officer, gave two reasons for the shift. The first is confidentiality. Some client matters involve information the firm considers too sensitive to hand to any outside party, cloud vendor or model provider included, and Mendoza said plainly that in those cases "we don't want to put it to any cloud vendor." The second reason is pricing risk. AI providers have been subsidizing token costs to win enterprise customers, and Mendoza wants Latham insulated from what happens once that subsidy narrows and usage-based billing takes over in earnest.


The timing is not incidental. Big Law is under steady pressure from corporate clients, particularly banks, to hold or lower fees even as the work involved shrinks with AI assistance, and a growing set of legal AI vendors is competing for the budget that pressure frees up. Latham's decision to build infrastructure rather than buy more vendor seats is a bet that owning the stack gives it more room to absorb that fee pressure over time than continuing to pay per-token or per-seat pricing to outside providers.


The Economics Driving the Decision


Latham's move looks unusual only if AI compute has to be rented by default. Across other industries, the same calculation is already playing out, and the pattern is consistent: ownership wins once usage is high and predictable enough that the GPUs rarely sit idle.

Bristol Myers Squibb runs Nvidia DGX SuperPOD infrastructure for computational drug discovery through Equinix Private AI, keeping cloud access available for overflow capacity, and reports 55 percent lower costs than its previous setup. Lockheed Martin has consolidated more than 30 AI models onto its own on-premises AI Factory, supporting roughly 7,000 engineers out of a workforce of about 122,000, and Nvidia says the system runs at full utilization with periods of overutilization as internal demand grows. BNY's Eliza AI platform, also built on Nvidia DGX SuperPOD hardware, now supports more than 17,000 users across over 40 applications in development. Texas A&M's research cluster, an academic rather than corporate deployment but a useful data point on utilization economics, runs nearly 760 GPUs at 95 to 98 percent capacity.


None of these organizations abandoned cloud entirely. What they did was separate steady, high-volume baseline workloads that run continuously from the unpredictable spikes and short-lived experiments that still make sense to rent. AWS currently prices Nvidia B200 GPU capacity at roughly $12.36 per accelerator hour and H100s at around $5.19 per hour in major U.S. regions. Run either continuously for a year and the bill climbs into six figures per chip.


Own the hardware instead, and the math flips once utilization is high enough that the upfront cost spreads across enough productive hours. Lenovo's own cloud-versus-on-prem analysis, a vendor estimate rather than an independent benchmark, claims some sustained inference deployments can break even against hyperscale cloud pricing in as little as six months and deliver up to a seventeen-fold cost advantage per million tokens compared with renting equivalent capacity as a service. Dell says it booked more than $130 billion in AI server orders over the past twelve months, a figure that reflects how many organizations are running this calculation and landing on the side of ownership. AMD's own hybrid-deployment modeling points to a similar range, estimating that splitting workloads roughly evenly between local and cloud infrastructure can cut three-year costs by 40 to 60 percent compared with a cloud-only approach, with fully local deployment reaching break-even in under two years under its assumptions. Vendor estimates like AMD's and Lenovo's should be read with the appropriate skepticism given their commercial interest in the on-prem narrative, but the direction they point in matches what Latham, Bristol Myers Squibb, Lockheed Martin, and BNY are independently reporting from their own deployments.


Where This Fits the Open-Prem Inflection Point


I have been tracking this exact convergence since April 2025 under the Open-Prem Inflection Point framework: the point where open-source model performance, falling hardware costs, and data sovereignty requirements line up closely enough that self-hosted AI deployment becomes the more rational default for organizations with steady, high-volume usage and sensitive data to protect. The April 2026 update to that framework documented nine or more open-weight model families operating at or near frontier performance, including DeepSeek V3.2 and GLM-5, and put self-hosted inference costs in the range of five to twenty cents per million tokens against three to fifteen dollars for the equivalent proprietary cloud API call. Organizations processing more than two million tokens a day, the framework estimated, could recover their hardware investment within six to twelve months.


Latham fits that profile closely. A firm generating $8.3 billion in annual revenue, running matter work across hundreds of active engagements, and now employing enough in-house technical staff to fine-tune models rather than just prompt them, is precisely the kind of organization the framework describes crossing the line from renting to owning. The two reasons Latham gave for the move, data too sensitive to expose and hedging against rising consumption-based pricing, are the same two factors the Open-Prem framework treats as the strongest pulls toward on-premises deployment, alongside raw usage volume.


What makes the Latham case notable is the industry it is happening in. Law firms have generally been buyers of AI tooling rather than builders of AI infrastructure, and professional services as a sector has been slower than pharma, defense, or finance to make this kind of capital commitment. A law firm now doing what Bristol Myers Squibb, Lockheed Martin, and BNY have already done suggests the calculation has stopped being specific to industries with obvious data-sovereignty mandates or research-scale compute needs. It has become a general enterprise decision available to any organization with enough steady AI usage to justify the upfront cost.


What Comes Next


Latham's approach carries real costs the firm now owns directly. Running GPU infrastructure means running cybersecurity, cooling, power, networking, and hardware refresh cycles internally rather than passing that responsibility to a vendor. Legal IT Insider's reporting flagged this trade-off explicitly, describing the setup as requiring substantial capital investment and specialist technical expertise in exchange for the control and cost flexibility Latham is buying. That trade-off is exactly what the utilization data from Lockheed Martin and Texas A&M suggests is worth paying, but only once usage is high enough to justify it. A firm running its GPUs at 30 percent utilization would be better off renting.


The practical signal for other enterprise leaders watching this is that the decision now deserves an actual calculation, rather than the reflexive assumption that cloud is always cheaper and simpler for every workload. CIOs weighing this need their own usage numbers: tokens processed daily, the sensitivity of the data involved, and how exposed they are to further pricing changes from their model providers. Firms without Latham's scale or technical headcount may find a hybrid path more realistic, running steady internal workloads on modest owned hardware while still renting cloud capacity for spikes and one-off projects, the same split Bristol Myers Squibb has already put into production. Latham ran its own calculation and landed on full ownership. Given how the same math has already played out in pharma, defense, and finance, more large professional-services firms are likely to run it too, even if most start smaller than a firm with $8.3 billion in annual revenue and a technology organization the size of Latham's behind it.

David Borish is the author of The Tony Hawk Paradox: When Video Games Predict Reality and writes The AI Spectator at davidborish.com


 
 

JOIN THE AI SPECTATOR MAILING LIST

CONTACT

Contacting You About:

Thanks for submitting!

New York, NY           

Db @DavidBorish.com           

  • LinkedIn
  • Instagram
  • Facebook
  • X
Back to top

© 2026 by David Borish IP, LLC, All Rights Reserved

bottom of page