I've been renting in New York for five years now. Every year my landlord raises the rent, and every year I pay it, because the alternative is owning, and owning is for people with a different relationship to money than mine. Renting is flexible. Renting is easy. Renting also means the building can change the rules on you whenever it wants, which is a thing I think about every time a new fee is added (pretty soon I’ll be paying for elevator access).
I bring this up because a bunch of software companies just had the same realization, except the thing they were renting was intelligence, mostly from Anthropic and OpenAI. And this summer, a surprising number of companies decided to buy instead of rent. Lucky them.

Bring back this man
TL;DR
Last week, an AI infrastructure company called Fireworks raised $1.5B at a $17.5B valuation. Two years ago, it was worth $552M. You've probably never heard of Fireworks, and honestly, that's fine. One stat buried in the announcement caught my eye: over 95% of the 40 trillion tokens they serve every day come from models specialized on their customers' own proprietary data.
Most of the AI running through that platform is companies serving their own custom models, tuned on their own data. Renting a frontier model from OpenAI or Anthropic and calling it a day is no longer the only way serious companies do this.
The announcements this summer:
June 18: Harvey, the legal AI company, announced its own legal foundation models, built by adapting open-source models so law firms can "own their own intelligence."
June 29: Wix's CEO published a manifesto titled "We built our own AI models." His reasons: better results, way lower cost, and the freedom to improve on his own schedule. When Wix owns the model, he says, it can act on user feedback "not quarterly, not tied to a major release cycle, but daily." In his words: "You're not waiting for a third party to prioritize your use case."
Earlier this year, Intercom declared "the age of vertical models is here" with a custom model behind its Fin support agent, and Shopify's fine-tuned model took over production traffic for Sidekick's automation builder.
The reason this is suddenly everywhere is that building your own model stopped being a moonshot idea. Open-source models (Qwen, Llama, DeepSeek) closed most of the quality gap with the frontier labs. And a technique called LoRA reduced the amount of work: instead of training a whole new model, you take an existing open model and attach a small add-on that teaches it your specific task. That add-on is your "custom model," and it's roughly the size of a large PowerPoint file. A fine-tuning run now costs tens to hundreds of dollars.
So companies did the math on their API bills and decided to build instead of rent.
SPONSORED BY
You bought something online with a credit card. You may have even used the “right” credit card.
But did you:
Activate offers on your credit cards
Use a shopping portal
Check for heavily discounted gift cards
Shoppers who use Savewise earn 7-10x more points & miles for things they’re already buying.
Savewise finds every way for you to earn rewards and puts them all in one place.
Specialists are beating the giants, sort of
The results, when this works, are a little absurd.
Shopify fine-tuned an open-source model to do one thing: turn a merchant's plain-English request (say, email me when inventory drops below 10) into a working automation. Their custom model is 68% cheaper than the frontier model it replaced and 2.2x faster, and Shopify says it now beats the big general models at this specific job. At launch, the fine-tune matched the old system on benchmarks, then performed 35% worse with real merchants. Only the weekly retraining loop on real merchant conversations has closed the gap.

LinkedIn built a model called EON-8B to power its recruiter agent, and one evaluation agent running on it makes nearly 90% of Hiring Assistant's AI calls, scoring candidates against job requirements all day long. It's 75x cheaper to serve than GPT-4 at that task.
Intercom trained its Fin Apex model on billions of interaction data points from its own resolved support tickets, and on its benchmarks, it now beats Claude Sonnet at customer service: higher resolution rate, 65% fewer hallucinations, faster responses. (Intercom's whole near-death-to-$3.6B-sale arc got its own issue last month, if you want the full story.) If you need help at a store, do you want a Nobel physicist or a retail assistant with 20 years of experience? You want the retail assistant (but you’d have a lot more fun with the physicist). Frontier models have to be good at everything. Your model only has to be good at your thing.
Cursor trained its own coding model, Composer, that generates about 4x faster than models of similar intelligence. They admit that GPT-5 is still smarter, but they bet that a fast model beats a brilliant model that you have to wait for. And Notion fine-tuned smaller models to cut search latency from 2 seconds to 350 milliseconds, with a line every product person should tattoo somewhere visible: "Latency is perceived as search quality."

There's also a reliability angle. Kyle Corbitt, who ran the fine-tuning startup OpenPipe, put numbers on it: a frontier model follows complicated conditional instructions maybe 75-80% of the time. A model fine-tuned on your task hits essentially 100%. The training run costs maybe $50 or $100.
So: cheaper, faster, better at the one job, more reliable, and it compounds weekly. Case closed, go train your model, right?
No. Please don't. Not yet, anyway.
Renting got too good to quit
The same two years that made building cheap also made renting nearly unbeatable.
Frontier models eliminated most of the need to fine-tune. Reasoning models now think through your weird edge cases at inference time, while they're answering, instead of needing your quirks trained into them beforehand. And with million-token context windows plus retrieval, you can hand a model your entire company brain in the prompt. In the past, you fine-tuned because the model couldn't hold your knowledge. Now it can just read it.
This spring, OpenAI shut down its own fine-tuning platform. The reason they gave developers in the wind-down email: prompt-based approaches are now cheaper and faster, and the newer models follow instructions well enough that fewer people need custom ones. When the biggest AI lab in the world exits the custom-model business because its base models got too good, believe them.
Many companies who’ve built models in the past have flopped. Bloomberg burned 1.3 million GPU-hours building BloombergGPT, a custom 50-billion-parameter finance model, in 2023 (Bloomberg never disclosed the cost; outside estimates put it in the millions). Within six weeks, researchers ran GPT-4 against the finance specialists, BloombergGPT included, and the generalist won most of the events. There was no BloombergGPT-2. Character.AI, after raising ungodly sums to train its own models, gave it up and moved to third-party models, explaining that "many more pre-trained models are now available." Woof.
And then the less fun stuff: evals. Before Cursor could train Composer, it had to build an entire internal benchmark to measure what "good" even means for its product. Hamel Husain, one of fine-tuning's loudest defenders, says it flat out: "It's impossible to fine-tune effectively without an eval system..." Before you can train your model, you first need to codify what is correct.
And the surveys back up the caution. Menlo's big enterprise study found fine-tuning is still niche; most companies get everything they need from prompting and retrieval on rented models. The industry has settled into a barbell, a split that analysts documented right when OpenAI pulled the plug: a small elite of AI-native companies training their own models on proprietary data, and everyone else renting frontier intelligence and building great harnesses around it. Both sides are right. They're just different companies.
SPONSORED BY
Built by yours truly → a 2-minute pre-game routine for high-stakes moments:
The important job interview in 5 minutes.
That big board presentation starting in 10 minutes.
The hard conversation you’ve been putting off with your team.
"Just breathe" doesn't work when you're already spiraling, and Calm and Headspace aren't built for the panic before the big moment.
Primo is. It walks you through a 2-minute routine using science-backed techniques from Stanford and Harvard.
The builders all had the same five things
Go back through every company in this piece, and the same pattern shows up.
They all had crushing volume on one narrow task. Shopify's automation requests. LinkedIn's billions of candidate scores. Fin's almost 2 million resolved issues a week. Nobody fine-tunes for a feature that users touch once.
They all had an API bill that had become its own problem. At that volume, frontier model pricing is $$$$. Shopify cut the cost of its feature by 68%; LinkedIn's swap took a dollar down to a penny. The savings matter as your usage goes up.
They all had proprietary data that no lab can get. Cursor's edit streams. Intercom's resolved tickets. Spotify's listening history. The model is how you cash in a data moat you already had (assuming you had one).
They all had latency as a product feature. Something users physically feel thousands of times a day.
They all had evals before they had models. You need the grading rubric before handing out the test.
And none of them started there. Every one of these companies ran on rented frontier models first, learned what their customers actually needed, accumulated the data, and built later.
The Test
You should build your own model only when you can answer yes to at least four out of five:
Volume. Do I have high, predictable traffic on one narrow, repeated task?
Cost. Is my API bill for that task big enough that cutting it by two-thirds would change my margins, my pricing, or my runway?
Data. Do I have proprietary interaction data (my users' edits, tickets, transactions) that no lab can buy or scrape?
Latency. Is speed something my users actually feel as product quality?
Evals. Do I already have an automated way to measure whether an answer is good?
Fail more than one, keep renting.
Score yourself right now. Most founders reading this will score a 1 or 2 out of 5. As you accumulate more data and burn through more tokens, you should reevaluate.
I’ve mentioned this before, but I think it’s exciting that regular old companies can build models that beat ones from AI labs. Pretty soon, your neighborhood car wash is going to have one to optimize its soap delivery.
If you’re a founder or working at a startup, I’d love to hear how you score.
Btw next week I’m on holiday, so there will be no newsletter. Try not to miss me too much.
— Amaraj (aka a lifelong renter, in every sense now)
The Meme



