AI & ML
Free AI API Tiers That Actually Last: 13 Options for 2026
Thirteen free AI and LLM API tiers that reset monthly or never expire, most with no credit card: what each gives you, the limit, and who it suits.
The free AI APIs worth building on are the ones with a rate limit or a refilling allowance, not a signup credit. Google's Gemini Developer API, Groq, Cohere and OpenRouter give you rate-limited access with no expiry date and no credit card; Mistral AI, IBM watsonx.ai and Hugging Face top up a free allowance every month. Start there if a side project has to stay free for longer than a trial.
Thirteen providers qualify. Every amount and limit below comes from the free AI API offers we track, checked against each provider's own pricing or rate-limit page on 19 September 2026.
What counts as a tier that lasts
One rule: the free part has to come back. A tier qualifies if it is ongoing — a rate limit with no end date — or if it refills monthly or daily. A one-time credit does not, however large: a project built on it has a shutdown date from the day you sign up. Those get a short section of their own, for contrast.
Two things are recorded for each tier: whether you need a card, and when the allowance resets. Twelve of the thirteen need no card. The exception is Vercel AI Gateway, which asks for a payment method before it activates. There are no quality or speed rankings here, and no benchmarks. Where a provider does not publish its cap, the entry says so.
Model APIs with a standing free tier
No balance to run down: access capped by a rate limit, for as long as the tier exists.
Google Gemini Developer API
The Gemini Developer API has an ongoing free tier: no expiration, no card, rate limits set per model. The ceiling depends on which model you call, so check the limit for yours before you design around it. It suits a side project making steady, modest calls.
Groq
Groq has an ongoing free tier with no expiration and no card. Its rate limits apply per organization, so the organization is the unit — extra keys inside it do not add headroom. It suits one project under one account.
Cohere
Cohere offers a free trial tier that is, despite the name, ongoing. Trial keys are rate-limited and — the important part — non-commercial. Fine for prototypes, internal tools and learning projects; the wrong base for anything you charge for.
OpenRouter
OpenRouter keeps a catalog of free models behind one API, capped at 20 requests per minute. The daily cap is not a fixed number: it depends on your account balance. No card is needed. It suits trying several models through the same interface before you commit to one.
NVIDIA NIM
NVIDIA NIM offers free hosted inference endpoints on build.nvidia.com, ongoing and without a card. The free access is credit-based, with no published figure we can quote, so treat it as an evaluation budget.
Replicate
Replicate lets you run a selection of models for free through its Try for Free collection, no card. The number of runs is limited, and the exact cap is not published. Good for a quick test of a model, not for a feature your users depend on.
Free credits that refill
A balance or a token budget, reset on a schedule — you can always see how much is left.
Mistral AI
Mistral AI lists $10 a month in API credits on its free plan, refreshed monthly, no card. The same plan limits messages, web searches and coding sessions. Once you know roughly what a request costs, you know how many you get.
IBM watsonx.ai
The watsonx.ai Lite plan gives you 300K tokens a month, reset monthly, no card. The catch is model choice: Lite has its own model selection, so confirm the one you want is on it. It suits anyone who would rather count tokens than dollars.
Hugging Face
Hugging Face gives free accounts $0.10 a month in inference credits, usable across its serverless inference providers and reset monthly. Ten cents is a testing budget, not a serving one. The more generous piece is ZeroGPU on Spaces: about five minutes of GPU a day on a free account, at medium queue priority, reset daily. That suits a public demo, not a backend. See Hugging Face.
Vercel AI Gateway
Vercel AI Gateway includes $5 every 30 days. It is the one tier here that needs a card: a payment method is required to activate it. It suits a project already on Vercel that wants model calls on the same bill. See Vercel.
Embeddings, search and scraping
Retrieval-style apps also need embeddings, reranking, search results and page content. Three providers cover those.
Jina AI
Jina AI gives free, rate-limited API access across its Reader, Embeddings and Reranker APIs: 500 requests per minute and 2M tokens per minute, per key. Ongoing, no card. For a retrieval side project, that is one account covering three steps of the pipeline.
Exa
Exa gives you a $20 credit at signup, then its free tier adds $10 in credits every month after that. No card. The signup credit is one-time; the monthly $10 is what makes it last. It suits an app that feeds search results to a model.
Firecrawl
Firecrawl includes 1,000 credits a month on its free plan, refreshed monthly and rate limited, no card. It suits pulling page content into a pipeline at side-project scale.
Side by side
| Provider | What is free | Card | Resets |
|---|---|---|---|
| Google Gemini Developer API | Rate-limited access, limits per model | No | Ongoing |
| Groq | Rate-limited access, limits per organization | No | Ongoing |
| Cohere | Rate-limited trial keys, non-commercial | No | Ongoing |
| OpenRouter | Free model catalog, 20 requests/min | No | Ongoing |
| NVIDIA NIM | Hosted inference endpoints, credit-based | No | Ongoing |
| Replicate | Runs on selected models, cap unpublished | No | Ongoing |
| Mistral AI | $10/month in API credits | No | Monthly |
| IBM watsonx.ai Lite | 300K tokens/month | No | Monthly |
| Hugging Face | $0.10/month credits; ~5 min/day ZeroGPU | No | Monthly; daily |
| Vercel AI Gateway | $5 every 30 days | Yes | Every 30 days |
| Jina AI | 500 requests/min and 2M tokens/min per key | No | Ongoing |
| Exa | $20 at signup, then $10/month | No | Monthly |
| Firecrawl | 1,000 credits/month | No | Monthly |
One-time credits, for contrast
Real money, but each has an end — good for an evaluation, not for running a side project on. All but Cerebras need no card.
- Claude API — one-time signup credits for new users, first-party Claude API only. See Anthropic API.
- Alibaba Cloud Model Studio — 1M tokens per model, valid 90 days from activation.
- Cerebras — a $5 trial credit for 30 days; card required.
- Scaleway — 1M tokens once, across its generative APIs, including 60 minutes of audio transcription.
- Voyage AI — 200M embedding and rerank tokens per account on current models (50M on legacy ones), excluding the Batch API.
- DeepInfra DeepStart — 1B inference tokens, for companies founded within the last two years that have raised $250K–$10M.
If you have raised money, the larger one-off grants are in startup programs: cloud and inference credits sized for companies rather than side projects. The guide to startup credits for cloud and AI sorts them by who actually qualifies.
How to stack free tiers without breaking the rules
Using two or three of these together is normal. Using them to fake capacity is not. The line is the account.
- One account per provider. A limit set per account or per organization is meant for one of you. Opening more accounts to multiply it is exactly what free-tier terms are written to stop, and the usual result is losing the tier.
- Fail over between providers, not between accounts. When your primary returns a rate-limit error, falling back to a different provider's free tier is ordinary engineering. Rotating the same traffic through several accounts at one provider is not.
- Put a thin interface in front. One function that takes a prompt and returns text, with the provider chosen by config. When a tier changes, switching costs an afternoon rather than a rewrite; expect to adjust prompts between models.
- Read the commercial-use terms before launch. The day a side project takes money, re-read the terms of every tier it touches.
- Count your own usage. Log requests and tokens per provider so you see a limit coming, instead of finding out from an error in production.
FAQ
Is there a free LLM API with no credit card?
Yes, several. Google's Gemini Developer API, Groq, Cohere, OpenRouter, Mistral AI and IBM watsonx.ai all have a recurring free tier that needs no card. Of the recurring tiers in this list, only Vercel AI Gateway asks for a payment method.
Which free AI API tiers never expire?
Gemini, Groq, Cohere, OpenRouter, NVIDIA NIM, Replicate and Jina AI are ongoing, with no end date. Mistral AI, IBM watsonx.ai, Hugging Face, Exa and Firecrawl refill monthly, Vercel every 30 days, and Hugging Face's ZeroGPU quota daily. Any of them can still change.
Can I use a free AI API tier in a commercial product?
It depends on the provider. Cohere's trial keys are explicitly non-commercial; for the rest, the answer is in each provider's terms, not in the size of the tier. Check before you charge anyone.
What is the difference between a free tier and free credits?
A free tier gives you rate-limited access with no balance: when you hit the limit, you wait. Free credits are a balance that runs down. Credits that refill monthly, like Mistral AI's or Exa's, behave like a tier; one-time credits, like Cerebras's $5 trial, are a countdown.
How current are these limits?
Every figure here was checked against the provider's own page on 19 September 2026. Free tiers change, so the AI API section of the offers page is where we keep them up to date.
The full list, with coding, media and startup offers alongside, is on the offers page. Voice has its own comparison of free text-to-speech and transcription APIs. For what to build with once you have a key, browse the AI & ML directory.