Each page here puts two model APIs or two serving stacks side by side on the same questions: what the publisher states about price, data use, retention, model lifecycle and exit, and what running the model yourself changes. Every fact links to its source and carries the date it was read.
| Page | What the page answers | Facts checked |
|---|---|---|
| OpenAI API vs Claude API: prices, data terms, retention and model retirement compared | OpenAI's API and the Claude API on what each publishes: token prices, batch and fast tiers, training use, retention, regions and retirement notice. | Sep 23, 2026 |
| OpenAI API vs a self-hosted LLM: what you pay, what you control, what can change under you | OpenAI API prices, data terms and retirement notice, read on 23 September 2026, against an open-weight model on your own GPUs: the break-even in full. | Sep 23, 2026 |
| vLLM vs Ollama: one machine for one user, or shared GPUs for many | Ollama is built for one user on one machine, vLLM for many users on shared GPUs. Defaults, licences, hardware and one published benchmark, compared. | Sep 23, 2026 |