Build vs Buy: LLM API vs Fine-Tuning Your Own
Almost everyone should start with an API. Here is the honest framework for when that stops being true.
Every few months a team decides their product needs "its own model." Usually they are staring at an API bill, or a competitor who claims to have trained something proprietary. It feels like the grown-up move. It almost never is. Start with an API. Stay on it longer than feels impressive. Train your own weights only when you hit a specific wall the API cannot climb. This is how to spot that wall.
#Why the API is the right default
An API call gives you a frontier model that a lab spent hundreds of millions of dollars training, for a few dollars per million tokens. Every upgrade lands for free. You write no training code, run no GPUs, debug no data pipeline. For most products the model was never the hard part. The hard part is the prompt, the retrieval, and the evaluation wrapped around it.
Most quality problems people blame on the base model are not base-model problems. Bad outputs come from a vague prompt, missing context, or no examples. Fine-tuning to paper over a weak prompt is like buying a faster car to fix your bad directions. Fix the directions first. Nine times out of ten the API was fine and the harness around it was thin.
#The walls that actually justify training
There are real reasons to own a model, and they are narrower than the hype suggests. Cost at volume: if you run millions of nearly identical calls a day, a small fine-tuned open model can be an order of magnitude cheaper per call, and that gap eventually pays for the engineers who run it. Hard constraints: if you need sub-100ms responses, or the model has to run on-prem or offline, you cannot call someone else's cloud. A task frontier models genuinely fail: a narrow domain with private data and a fixed output format, where you have thousands of high-quality examples the base model has never seen.
What is not on the list: "we want it to sound like our brand," or "we have proprietary data." Those are prompt and retrieval problems. You can put proprietary data in the context window at request time. You do not need to bake it into weights to use it.
#There is a middle option, and it is usually the answer
Build-vs-buy is a false binary. The real ladder has rungs. Start with a good prompt on an API. Add few-shot examples. Add retrieval so the model reads your data at request time. Reach for fine-tuning only after those are exhausted, and even then the first stop is a hosted fine-tune through the same provider, not your own GPUs.
Run this test before you train anything. Take 50 of your hardest real cases and write down the answers you actually want. Run your best API prompt against them and score it. If you are failing on format or a narrow style, fine-tuning will help. If you are failing on reasoning or facts, it will not, and a smaller model you host yourself will be worse. Fine-tuning teaches a model a shape, not new intelligence. If you cannot name the shape you are teaching, you are not ready to train.
#The bottom line
Owning a model is a real commitment. Data pipelines, eval harnesses, GPU ops, and re-training every time the open-source frontier moves. Sometimes that cost is worth it, and when it is, the case is obvious and measurable. If you cannot name the exact wall the API hit and show the failing test cases, you have not hit it yet. Keep calling the API and go fix your prompt.