We hadn't done a full audit of this directory in a while, and it showed. Over the last few months a popular free tier was shut down, another one quietly started asking for a payment method, and every big provider retired at least one model that tutorials (including some of ours) still recommend. This is the result of going through the whole list, provider by provider, against official docs and pricing pages on October 4, 2026.
A note on method before the findings. Whenever possible I used the provider's own documentation, deprecation page or public model list, not a third-party "best free APIs" roundup. Where I could only find third-party evidence, or where the provider's console needs a login I don't have, I say so in the text instead of presenting it as fact. Everything here is a snapshot: free tiers change monthly, and the directory on the home page and the changelog are the living version of this article.
The short version
- Gone: GitHub Models (retired July 30, 2026) and Hyperbolic's serverless inference API (retired; only GPU rental remains).
- No longer free in practice: Chutes.ai (since 2025), Together.AI (only one model at $0), Novita AI (no $0 models left, trial credit only), and Cerebras' no-card tier, which now needs a verified payment method.
- Retired models you should stop hard-coding: Gemini 2.0 and 1.5, Groq's
llama-3.3-70b-versatileandllama-3.1-8b-instant, Groq Compound, Mistral'sopen-mistral-nemo, and most of Cerebras' Llama and Qwen3 32B line-up. - New and worth using: Gemini 3.5 to 3.8 Flash, GPT OSS 120B almost everywhere, Qwen3.8 27B, the new OpenRouter free models (Nemotron 3 Ultra, Laguna, Inkling), Hetzner's free inference API, and a new wave of free-tier gateways.
- The pattern: "free forever, no card" is shrinking. "Free credits with a card on file" and "free models behind a gateway" are growing. Plan for it.
Part 1: Providers that shut down or lost their free tier
GitHub Models: retired
GitHub Models was one of the easiest ways to get free access to frontier-class models with nothing but a GitHub account. GitHub fully retired it on July 30, 2026, according to its official changelog. If you still have it in a config file, requests will fail. It stays on this site as a discontinued entry so old links and old tutorials still resolve to something that explains what happened.
Hyperbolic: serverless API retired
Hyperbolic used to be a staple of "free credits" lists ($1 on signup, plus Llama 3.1 405B and DeepSeek models). Their documentation now states plainly that the serverless inference API and the playground have been retired, that the old models are no longer available, and that requests to them no longer work even with a valid key. Existing credits can still be used for GPU rentals and storage. We have marked Hyperbolic as discontinued; our ultimate guide now carries the same note.
Cerebras: the no-card tier ended
Cerebras is still extremely fast and still has a free tier, but it is no longer a "sign up and go" one. Accounts get a $5 free credit that expires 30 days after it is granted, and both the playground and the API stay inactive until a verified payment method is added. The free trial limits are 5 requests per minute, 30,000 uncached and 90,000 total tokens per minute, and 1,000,000 tokens per hour and per day. That is still good for prototypes; it is just no longer something you can hand to a student without a card.
Together.AI, Novita, Chutes, Venice: free in name only
- Together.AI: its public model list marks exactly one chat model as free,
Prism-ML/Ternary-Bonsai-27B. The four previously free models we listed (Llama 4 Scout, DeepSeek R1 Fast, two Apriel thinkers) are no longer $0. - Novita AI: its own comparison article now says no model in the current catalog is priced at $0 per token. What remains is a $0.50 trial credit (valid one year) and a $10-for-both-sides referral program capped at $500.
- Chutes.ai: no free tier for new signups since 2025; the cheapest plan is $10/month.
- Venice.ai: its pricing page lists API access on the Free plan as "pay with credits" (100 credits = $1). We kept the listing but flagged it, because the line between "free account" and "free API" is exactly the kind of thing that wastes an afternoon.
Other changes worth knowing
- Google AI Studio still has a generous free tier, but only for Flash, Flash-Lite, Gemma and a few specialty models. Gemini 3.1 Pro Preview has no free tier on the official pricing page.
- Nebius Token Factory asks for a bank card at onboarding.
- SiliconFlow keeps a rotating set of permanently free small models (Qwen3-8B and DeepSeek-R1-Distill-Qwen-7B at the time of writing). Third-party sources report mandatory real-name verification since May 2026, which we could not confirm without an account.
- ModelScope offers 2,000 free API calls per day with a 500-call cap per model, but requires an Alibaba Cloud account binding. That is real friction outside China.
Part 2: Models that were retired (and what to use instead)
The quieter breakage is model IDs. A provider can keep its free tier and still kill the specific model your app calls. These are the retirements we confirmed against official deprecation pages:
| Provider | Retired | Use instead |
|---|---|---|
gemini-2.0-flash, gemini-2.0-flash-lite (June 1, 2026); 1.5 family already gone |
gemini-3.8-flash, gemini-3.5-flash-lite |
|
| Groq | llama-3.1-8b-instant, llama-3.3-70b-versatile (Aug 16); qwen/qwen3-32b, llama-4-scout (Jul 17); qwen3.6-27b (Sep 14); groq/compound and compound-mini (Sep 21) |
openai/gpt-oss-120b, openai/gpt-oss-20b, qwen/qwen3.8-27b |
| Cerebras | llama3.1-8b and qwen-3-235b-a22b-instruct-2507 (May 27), qwen-3-32b and llama-3.3-70b (Feb 16), llama-4-scout, zai-glm-4.7 (Aug 17) |
gpt-oss-120b, qwen-3.8-27b |
| Mistral | open-mistral-nemo (Jul 31), mistral-medium-2508 (Aug 31), open-mistral-7b and open-mixtral-8x7b (long gone) |
mistral-medium-2604, mistral-small-2603, mistral-large-2512 |
| DeepSeek | deepseek-v4-flash is now a legacy alias |
deepseek-flash (V4.1-Flash), deepseek-v4-pro |
| Scaleway | llama-3.1-70b-instruct (May 2025), gemma-3-27b-it (Aug 1, 2026) |
llama-3.3-70b-instruct, gemma-4-26b-a4b-it |
| Nebius | Llama 3.3 70B, Qwen3 32B and Hermes 4 70B (Aug 31); eleven models removed June 22 | Nemotron 3.5 Lightning, MiniMax M3, Qwen3.5 397B |
Two themes stand out. First, GPT OSS 120B has become the default replacement everywhere: Groq, Cerebras, SambaNova, Cloudflare, OVHcloud and Nebius all list it. If you need one model that is available on many providers, that is currently the safest bet for portability. Second, Llama 3.x is being phased out of free tiers. A lot of tutorials still point at Llama 3.1 and 3.3 model IDs; many of those will return a "model not found" error today.
Part 3: What's new and actually free
OpenRouter's free catalog was refreshed almost entirely
Of the 16 free models we previously listed on OpenRouter, only one (Nemotron 3 Ultra, under a new model ID) is still on the official free collection page. The current list includes NVIDIA Nemotron 3 Ultra (550B total, 1M context), Nemotron 3 Super and 3.5 Lightning, Poolside Laguna S 2.1 and XS 2.1, Qwen3.8 27B, Thinking Machines Inkling and Inkling Small, Cohere North Mini Code, Dots3-Note Preview and Apodex 1.1 Mini. One caveat that matters: the Laguna listings state that, on the free route, your inputs and outputs may be used to train the provider's models. Do not send private data to a free model without reading that line. Request limits are also modest: third-party testing reports 50 requests per day on a free account, rising to 1,000 per day after $10 of lifetime credit.
Google's Gemini 3.5 to 3.8 Flash
Google's pricing page lists free-tier access for gemini-3.8-flash,
gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash,
gemini-3.5-flash-lite, gemini-3.1-flash-lite, the 2.5 Pro / Flash /
Flash-Lite models, and Gemma 4. Per-model limits are shown per project in AI Studio rather than in
the docs, so check yours before you plan around a number. Also remember that outside the UK, EEA and
Switzerland, free-tier prompts can be used to improve Google's products.
Hetzner Inference API: a new free option from an EU host
Hetzner launched an experimental OpenAI-compatible inference API in July 2026. It needs only a free
Hetzner account and a token, no card, and currently serves a single model,
Qwen/Qwen3.6-35B-A3B-FP8 (35B MoE with about 3B active parameters, 262K context,
vision). The endpoint is https://inference.hetzner.com/api/v1 and the per-key limits are
generous: 3M input and 60K output tokens per minute, 500M input and 5M output tokens per day. The
catch is in the word "experimental": no SLA, no performance guarantee, and Hetzner says it will email
users before any billing starts. It is great for prototypes and European data-residency
experiments, not for production.
Cloudflare Workers AI: a bigger free catalog
The allowance is unchanged at 10,000 neurons per day, but the catalog behind it has moved on:
gpt-oss-120b and gpt-oss-20b, Llama 4 Scout, Gemma 4 26B,
Qwen3 30B-A3B, Mistral Small 3.1 and GLM 4.7 Flash are all available. Some frontier models
(Kimi K2.6 and K2.7-code, GLM 5.2 and 5.3, DeepSeek variants) require a paid billing method and
cannot be used on the free allocation.
Free models you can use today, in one table
| Provider | What's free | Watch out for |
|---|---|---|
| Google AI Studio | Gemini 3.x Flash and Flash-Lite, 2.5 models, Gemma 4 | Pro models paid; prompts may be used for training outside EEA/UK/CH |
| Groq | GPT OSS 120B / 20B, Qwen3.8 27B (preview), Whisper | Preview models can be discontinued with little notice |
| OpenRouter | ~14 :free models, 1M-context options |
Catalog rotates; some free routes train on your prompts; 50 req/day |
| Cloudflare Workers AI | 10,000 neurons/day on open models | Big models need paid billing |
| Hetzner Inference | Qwen3.6-35B-A3B, very high limits | Experimental, no SLA |
| Z.AI (GLM) | GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash (vision) | Free-model rate limits not published |
| Mistral | Free API tier on current models (limits shown in the console) | Retired IDs return errors; use the dated IDs above |
| Cohere | Trial key, 20 req/min on chat models | 1,000 calls/month cap, non-commercial |
| Requesty | A handful of free routed models | 50 requests/day for new orgs (200 for paying orgs), shared across free models |
A new wave of free-tier gateways
Several gateways now give you a free allocation as a way to get you into their platform. We have added Kilo AI Gateway (a rotating set of default free models), Vercel AI Gateway (a monthly free credit usable on a subset of models, which stops applying as soon as you buy credits), Api.Airforce (1 request per minute, 1,000 per day) and Routeway (a 200-requests-per-day starter plan on experimental models). They are convenient, but they are also exactly where the "free" line is easiest to cross by accident, so treat the details as partly verified and read the provider's page before depending on them. We also reviewed and declined several submissions that could not show verifiable limits, and one that ran behind a temporary tunnel URL.
Part 4: Free, trial, and "free to start" are different things
The biggest source of confusion we see in reports from visitors is not wrong numbers, it is wrong categories. It helps to sort every offer into one of four buckets:
- Permanent free tier: a rate-limited allowance that renews and has no expiry (Google AI Studio, Cloudflare, Groq, Hetzner for now).
- Renewable credits: free capacity that resets on a schedule but is capped by spend or request count (OpenRouter, Vercel AI Gateway, Hugging Face's $0.10 monthly routing credit).
- One-time trial credits: a balance that runs out or expires: Cerebras ($5, 30 days), Scaleway (1M tokens), Alibaba Model Studio (1M tokens per model for 90 days), Fireworks ($1), Novita ($0.50).
- "Free to start": an account is free but API use is paid, as on Venice's current pricing page. This is not a free tier.
Two small clarifications that matter in practice. Fireworks' 10 requests per minute applies to accounts without a payment method, and the 6,000 RPM ceiling only applies with a card and active credits. And a free Hugging Face account gets only $0.10 in monthly Inference Providers credit ($2 on PRO), which is a testing allowance, not an API plan.
Part 5: How to build so the next shutdown doesn't hurt
- Never hard-code a single model ID. Put the model name in config, not in code.
Several providers rename models (DeepSeek's V4 Flash became
deepseek-flash; Groq's Qwen 3.6 became 3.8) and the old name only works for a while. - List models at startup. Most OpenAI-compatible providers expose
GET /v1/models. If your configured ID is missing, log it loudly and fall back instead of failing at the first user request. - Use a fallback chain. We wrote the pattern up in Never Hit a Rate Limit Again. The same chain also protects you from a provider disappearing, not just from 429s.
- Prefer portable models. Choosing GPT OSS 120B or a Qwen3.8 variant gives you the widest set of providers to fall back to.
- Treat preview models as temporary. Groq's own docs call its preview models "evaluation only"; several have been removed within weeks. Do not build a product on one.
- Mind the data policy. Free routes often come with training clauses. Keep private data on paid or self-hosted paths.
- Watch for changes instead of discovering them. Our changelog (and its RSS feed) logs every model added or removed, and the setup generator always uses the current list.
What we couldn't verify
An honest audit includes its gaps. We could not confirm the free-model lists inside the consoles of SiliconFlow and ModelScope (they require a login), the exact free allowance of Aion Labs (third parties report 15 requests per minute and 20K tokens per day), or whether Inference.net still offers free LLM inference, since its site now emphasizes tracing and gateway products. We also found no evidence that xAI still offers free API credits, so Grok models are flagged accordingly. These entries carry an explicit "verify" note on their pages instead of a confident number.
Wrap up
The free LLM landscape in October 2026 is healthier in one way and more fragile in another. There are more free models than ever, and the quality of the free ones is genuinely good: a free Gemini 3.x Flash, GPT OSS 120B or Qwen3.8 27B would have been remarkable a year ago. But the providers' incentives have shifted from "attract everyone with no card" to "get you onto the platform first", and the model catalog under every free tier changes faster than any tutorial. Build for that: configurable models, a fallback chain, and a habit of checking the changelog. If you spot something we got wrong, the Report an Issue button on every provider page goes straight to the moderation queue.
Related reading: the ultimate free LLM API guide, OpenRouter alternatives, and how to use OpenRouter's free models.