Awesome Free LLM APIs

Awesome

LLM APIs with permanent free tiers for text inference.

All endpoints are OpenAI SDK-compatible unless noted. Each link points to the provider's API key page.


Manifest

All of those free LLM APIs are available at manifest.build - make reliable agents.

Contents

Provider APIs

APIs run by the companies that train or fine-tune the models themselves.

Aion Labs ๐Ÿ‡ฎ๐Ÿ‡ฑ

Permanent free tier, no credit card required. 15 RPM, 20K tokens/day. Specialized for roleplay and storytelling.

Base URL: https://api.aionlabs.ai/v1

Model NameContextMax OutputModalityRate Limit
Aion 2.5128K32KText (roleplay)15 RPM, 20K TPD
Aion 2.0128K32KText (roleplay)15 RPM, 20K TPD
Aion-RP 1.0 (8B)32K32KText (roleplay)15 RPM, 20K TPD
Aion 3.0128K32KText (roleplay, reasoning)15 RPM, 20K TPD
Aion 3.0 Mini128K32KText (roleplay, reasoning)15 RPM, 20K TPD

Cohere ๐Ÿ‡จ๐Ÿ‡ฆ

Free "Trial" API key, no credit card. 1,000 API calls/month. Non-commercial use only.

Base URL: https://api.cohere.com/v2

Model NameContextMax OutputModalityRate Limit
Command A+ (218B)436K64KText + Image20 RPM
Command A (111B)288K8KText20 RPM
Command R+128K4KText20 RPM
Command R128K4KText20 RPM
Command R7B128K4KText20 RPM
Command A Reasoning288K~4KText (reasoning)20 RPM
Command A Translate~9K~4KText20 RPM
Command A Vision128K~4KText + Image20 RPM
Command R7B Arabic128K~4KText20 RPM
Aya Expanse 32B128K~4KText20 RPM
Aya Vision 32B16K~4KText + Image20 RPM

Google Gemini ๐Ÿ‡บ๐Ÿ‡ธ

Free tier unavailable in EU/UK/Switzerland. Free-tier prompts may be used by Google to improve products. 1

Base URL: https://generativelanguage.googleapis.com/v1beta

Model NameContextMax OutputModalityRate Limit
Gemini 3.6 Flash1M65KText + Image + Audio + Video15 RPM, 1,500 RPD
Gemini 3.5 Flash1M65KText + Image + Audio + Video15 RPM, 1,500 RPD
Gemini 3.5 Flash-Lite1M65KText + Image + Audio + Video30 RPM, 1,500 RPD
Gemini 3.1 Flash-Lite1M65KText + Image + Audio + Video30 RPM, 1,500 RPD
Gemini 2.5 Flash1M65KText + Image + Audio + Video15 RPM, 1,500 RPD
Gemini 2.5 Flash-Lite1M65KText + Image + Audio + Video30 RPM, 1,500 RPD
Gemini 2.5 Pro1M65KText + Image + Audio + Video5 RPM, 50 RPD

Mistral AI ๐Ÿ‡ซ๐Ÿ‡ท

Free "Experiment" plan, no credit card. ~1B tokens/month. Prompts may be used to improve models.

Base URL: https://api.mistral.ai/v1

Model NameContextMax OutputModalityRate Limit
Mistral Medium 3.5 (128B)256K256KText + Image + Code~1 RPS, 500K TPM
Mistral Small 4256K256KText + Image + Code~1 RPS, 500K TPM
Mistral Large 3256K256KText~1 RPS, 500K TPM
Ministral 8B256K256KText~1 RPS, 500K TPM
Codestral256K256KCode~1 RPS, 500K TPM
Ministral 3B128K128KText~1 RPS, 500K TPM
Ministral 14B256K256KText~1 RPS, 500K TPM

Z AI (Zhipu AI) ๐Ÿ‡จ๐Ÿ‡ณ

Permanent free models, no credit card required.

Base URL: https://open.bigmodel.cn/api/paas/v4

Model NameContextMax OutputModalityRate Limit
GLM-4.7-Flash200K128KText (reasoning)1 concurrent request
GLM-4.5-Flash128K~96KText (reasoning)1 concurrent request
GLM-4.6V-Flash128K~4KText + Image1 concurrent request

Inference providers

Third-party platforms that host open-weight models from various sources.

Cerebras ๐Ÿ‡บ๐Ÿ‡ธ

Free tier with payment method required. Ultra-fast inference. 1M tokens/day cap. 64K context on free tier.

Base URL: https://api.cerebras.ai/v1

Model NameContextMax OutputModalityRate Limit
gpt-oss-120b131K (65K on free)32K (free) / 40K (paid)Text5 RPM, 30K TPM, 1M TPD
zai-glm-4.7 (deprecated Aug 2026)131K (64K on free)40KText5 RPM, 30K TPM, 1M TPD
gemma-4-31b131K (65K on free)32K (free) / 40K (paid)Text + Image15 RPM, 30K TPM, 1M TPD

Cloudflare Workers AI ๐Ÿ‡บ๐Ÿ‡ธ

10,000 Neurons/day free. 50+ models available on free tier.

Base URL: https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run

Model NameContextMax OutputModalityRate Limit
@cf/meta/llama-3.3-70b-instruct-fp8-fast131KShared w/ contextText10K neurons/day (shared)
@cf/meta/llama-4-scout-17b-16e-instructUp to 10MShared w/ contextMultimodal10K neurons/day (shared)
@cf/openai/gpt-oss-120b128KShared w/ contextText10K neurons/day (shared)
@cf/moonshotai/kimi-k2.7-code262KShared w/ contextText (code)10K neurons/day (shared)
@cf/google/gemma-4-26b-a4b-it256KShared w/ contextText10K neurons/day (shared)
@cf/zhipuai/glm-4.7-flash131KShared w/ contextText10K neurons/day (shared)
@cf/mistralai/mistral-small-3.1-24b-instruct128KShared w/ contextText10K neurons/day (shared)
@cf/deepseek-ai/deepseek-r1-distill-qwen-32b32KShared w/ contextText (reasoning)10K neurons/day (shared)
+ 42 more modelsVariesVariesText, Image, Audio, Embeddings10K neurons/day (shared)

GitHub Models ๐Ÿ‡บ๐Ÿ‡ธ

Free prototyping for all GitHub users. 45+ models. Per-request limits (8K in / 4K out).

Base URL: https://models.github.ai/inference

Model NameContextMax OutputModalityRate Limit
gpt-5200K32KText10 RPM, 50 RPD
gpt-4.11M32KText10 RPM, 50 RPD
gpt-4.1-mini1M32KText15 RPM, 150 RPD
gpt-4o128K16KText + Vision10 RPM, 50 RPD
o4-mini200K100KText (reasoning)10 RPM, 50 RPD
Llama-4-Scout-17B-16E-Instruct512K~4KText + Vision15 RPM, 150 RPD
Llama-4-Maverick-17B-128E-Instruct-FP8256K~4KText + Vision10 RPM, 50 RPD
Llama-3.3-70B-Instruct131K~4KText15 RPM, 150 RPD
DeepSeek-R164K8KText (reasoning)15 RPM, 150 RPD
Mistral-Small-3.1128K~4KText + Vision15 RPM, 150 RPD
+ 35 more modelsVariesVariesText / ImageVaries by tier

Groq ๐Ÿ‡บ๐Ÿ‡ธ

Free tier, no credit card. Ultra-fast LPU inference. 2

Base URL: https://api.groq.com/openai/v1

Model NameContextMax OutputModalityRate Limit
llama-3.3-70b-versatile131K32KText30 RPM, 1,000 RPD
llama-3.1-8b-instant131K131KText30 RPM, 14,400 RPD
openai/gpt-oss-120b131K65KText30 RPM, 1,000 RPD
openai/gpt-oss-20b131K65KText30 RPM, 1,000 RPD
groq/compound131K8KText30 RPM, 250 RPD
groq/compound-mini131K8KText30 RPM, 250 RPD
qwen/qwen3.6-27b131K16KText30 RPM, 1,000 RPD

Hugging Face ๐Ÿ‡บ๐Ÿ‡ธ

$0.10/month in Inference Provider credits for free users (subject to change). Routes to Fireworks, Together, Hyperbolic, Nebius, Novita, DeepInfra and others. Thousands of models.

Base URL: https://router.huggingface.co/v1

Model NameContextMax OutputModalityRate Limit
Meta-Llama-3.1-8B-Instruct128K~4KTextCredit-metered
gemma-3-4b-it131K~4KTextCredit-metered
phi-416K~4KTextCredit-metered
Qwen2.5-Coder-7B-Instruct131K~4KTextCredit-metered
Qwen2.5-7B-Instruct131K~4KTextCredit-metered
+ thousands of community modelsVariesVariesText, Image, Audio, Embeddings100K credits/month free

Kilo Code ๐Ÿ‡บ๐Ÿ‡ธ

Free models with no credit card required. kilo-auto/free auto-router dynamically routes to models in the free pool. 3

Base URL: https://api.kilo.ai/api/gateway

Model NameContextMax OutputModalityRate Limit
nvidia/nemotron-3-ultra-550b-a55b:free1M65KText~200 req/hr
stepfun/step-3.7-flash:free262K262KText~200 req/hr
nvidia/nemotron-3-super-120b-a12b:free262K262KText~200 req/hr
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free256K65KText (reasoning)~200 req/hr
inclusionai/ling-3.0-flash:free262K32KText~200 req/hr
poolside/laguna-s-2.1:free262K32KText (code)~200 req/hr
poolside/laguna-xs-2.1:free262K32KText (code)~200 req/hr
cohere/north-mini-code:free256K64KText (code)~200 req/hr
openrouter/freeVariesVariesText~200 req/hr

LLM7.io ๐Ÿ‡ฌ๐Ÿ‡ง

Zero-friction API gateway. No registration needed for basic access. 30+ models. GDPR-compliant.

Base URL: https://api.llm7.io/v1

Model NameContextMax OutputModalityRate Limit
deepseek-r1-0528โ€”โ€”Text (reasoning)30 RPM (120 with token)
deepseek-v3-0324โ€”โ€”Text30 RPM (120 with token)
gemini-2.5-flash-liteโ€”โ€”Text + Vision30 RPM (120 with token)
gpt-4o-miniโ€”โ€”Text + Vision30 RPM (120 with token)
mistral-small-3.1-24b32Kโ€”Text30 RPM (120 with token)
qwen2.5-coder-32bโ€”โ€”Text (code)30 RPM (120 with token)
+ ~24 more modelsVariesVariesText30 RPM (120 with token)

ModelScope ๐Ÿ‡จ๐Ÿ‡ณ

Free API-Inference for registered users. Requires Alibaba Cloud account binding + real-name verification. 4

Base URL: https://api-inference.modelscope.cn/v1

Model NameContextMax OutputModalityRate Limit
Qwen/Qwen3.5-35B-A3Bโ€”โ€”Text2,000 RPD total; <=500 RPD/model (dynamic)
Qwen/Qwen3.5-27Bโ€”โ€”Text2,000 RPD total; <=500 RPD/model (dynamic)
+ API-Inference-enabled modelsVariesVariesLLM, MLLMDynamic quotas + dynamic concurrency

NVIDIA NIM ๐Ÿ‡บ๐Ÿ‡ธ

Free with NVIDIA Developer Program membership. 100+ models. Rate-limited (no daily token cap).

Base URL: https://integrate.api.nvidia.com/v1

Model NameContextMax OutputModalityRate Limit
deepseek-ai/deepseek-v4-flash1M~64KText~40 RPM
nvidia/nemotron-3-super-120b-a12b262K262KText~40 RPM
nvidia/nemotron-3-nano-30b-a3b128K32KText~40 RPM
nvidia/llama-3.1-nemotron-ultra-253b-v1128K4KText~40 RPM
meta/llama-3.3-70b-instruct128K4KText~40 RPM
mistralai/mistral-nemotron128K8KText~40 RPM
google/gemma-4-31b-it128K8KText~40 RPM
mistralai/mistral-large-2-instruct128K4KText~40 RPM
minimaxai/minimax-m31M~64KText~40 RPM
mistralai/mistral-medium-3.5-128b262K262KText~40 RPM
nvidia/nemotron-3-ultra-550b-a55b262K262KText~40 RPM
openai/gpt-oss-120b131K131KText~40 RPM
openai/gpt-oss-20b131K131KText~40 RPM
deepseek-ai/deepseek-v4-pro128K~64KText~40 RPM
+ 85 more modelsVariesVariesText, Image, Video, Speech, Embeddings~40 RPM

Ollama Cloud ๐Ÿ‡บ๐Ÿ‡ธ

Free tier with qualitative usage limits. 400+ models from Ollama library. Not OpenAI SDK-compatible; uses Ollama API. 5

Base URL: https://api.ollama.com

Model NameContextMax OutputModalityRate Limit
deepseek-v4-pro128KModel-dependentTextSession/weekly limits (unpublished)
deepseek-v4-flash1MModel-dependentTextSession/weekly limits (unpublished)
minimax-m31MModel-dependentTextSession/weekly limits (unpublished)
kimi-k3128KModel-dependentTextSession/weekly limits (unpublished)
gpt-oss:120b128KModel-dependentTextSession/weekly limits (unpublished)
gpt-oss:20b131KModel-dependentTextSession/weekly limits (unpublished)
nemotron-3-ultra262KModel-dependentTextSession/weekly limits (unpublished)
mistral-large-3:675b128KModel-dependentTextSession/weekly limits (unpublished)
qwen3.5:397b131KModel-dependentTextSession/weekly limits (unpublished)
+ 10 more cloud modelsVariesVariesTextSession/weekly limits (unpublished)

OpenRouter ๐Ÿ‡บ๐Ÿ‡ธ

~22 free models (marked with :free suffix). OpenAI SDK-compatible. 6

Base URL: https://openrouter.ai/api/v1

Model NameContextMax OutputModalityRate Limit
nvidia/nemotron-3-super-120b-a12b:free262K262KText20 RPM, 50 RPD
openai/gpt-oss-20b:free131K32KText20 RPM, 50 RPD
cohere/north-mini-code:free256K64KText (code)20 RPM, 50 RPD
google/gemma-4-26b-a4b-it:free262K32KText + Image20 RPM, 50 RPD
google/gemma-4-31b-it:free262K32KText + Image20 RPM, 50 RPD
inclusionai/ling-3.0-flash:free262K32KText20 RPM, 50 RPD
nvidia/nemotron-3-nano-30b-a3b:free256Kโ€”Text20 RPM, 50 RPD
nvidia/nemotron-nano-9b-v2:free128Kโ€”Text20 RPM, 50 RPD
nvidia/nemotron-nano-12b-v2-vl:free128K128KText + Image20 RPM, 50 RPD
poolside/laguna-s-2.1:free262K32KText (code)20 RPM, 50 RPD
poolside/laguna-xs-2.1:free262K32KText (code)20 RPM, 50 RPD
+ ~12 more free modelsVariesVariesText / Image20 RPM, 50 RPD

OVHcloud AI Endpoints ๐Ÿ‡ซ๐Ÿ‡ท

Free anonymous tier (no API key, no signup): 2 RPM per IP per model. 20+ open-weight models hosted in EU. OpenAI SDK-compatible. 7

Base URL: https://oai.endpoints.kepler.ai.cloud.ovh.net/v1

Model NameContextMax OutputModalityRate Limit
Qwen3.5-397B-A17B131K~32KText2 RPM (anonymous)
gpt-oss-120b128K~32KText2 RPM (anonymous)
gpt-oss-20b128K~8KText2 RPM (anonymous)
Meta-Llama-3_3-70B-Instruct131K~4KText2 RPM (anonymous)
Qwen3.6-27B131K~32KText2 RPM (anonymous)
Qwen3.5-9B131K~8KText2 RPM (anonymous)
Qwen3-32B131K~32KText2 RPM (anonymous)
Qwen3-Coder-30B-A3B-Instruct262K~32KText (code)2 RPM (anonymous)
Qwen2.5-VL-72B-Instruct128K~8KText + Vision2 RPM (anonymous)
Mistral-Small-3.2-24B-Instruct128K~4KText2 RPM (anonymous)
Mistral-Nemo-Instruct-2407128K~4KText2 RPM (anonymous)
Mistral-7B-Instruct-v0.332K~4KText2 RPM (anonymous)

SambaNova ๐Ÿ‡บ๐Ÿ‡ธ

Free tier, no credit card. Ultra-fast RDU inference. 20 RPM, 200K tokens/day. 8

Base URL: https://api.sambanova.ai/v1

Model NameContextMax OutputModalityRate Limit
DeepSeek-V3.1128K~8KText20 RPM, 20 RPD, 200K TPD
DeepSeek-V3.2 (Preview)128K~8KText20 RPM, 20 RPD, 200K TPD
Meta-Llama-3.3-70B-Instruct128K~3KText20 RPM, 20 RPD, 200K TPD
gpt-oss-120b128K~128KText20 RPM, 20 RPD, 200K TPD
MiniMax-M2.7128K~192KText20 RPM, 20 RPD, 200K TPD
gemma-4-31B-it (Preview)128K~128KText + Image + Video20 RPM, 20 RPD, 200K TPD

SiliconFlow ๐Ÿ‡จ๐Ÿ‡ณ

Permanently free models, no credit card required. 200+ paid models also available.

Base URL: https://api.siliconflow.cn/v1

Model NameContextMax OutputModalityRate Limit
Qwen/Qwen3-8B131K131KText30 RPM, 60K TPM
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B131KConfigurableText (reasoning)30 RPM, 60K TPM

Glossary

AbbreviationMeaning
RPMRequests per minute
RPDRequests per day
TPMTokens per minute
TPDTokens per day
RPSRequests per second

Contributing

Know a free tier that's missing? Open a PR. Include the provider, endpoint, rate limits (link to their docs), and a few notable models. Trial credits and time-limited promos don't count.

Footnotes

  1. Free tier not available in the EU, UK, or Switzerland (available regions). โ†ฉ

  2. Groq rate limits were reduced in 2026. Most models now get 1,000 RPD on the free tier (down from 14,400). Llama 4 Maverick has been deprecated. See rate limits. โ†ฉ

  3. Kilo Code free model list changes frequently. nvidia/nemotron-3-super-120b-a12b:free is for trial use only โ€” prompts are logged by NVIDIA. Auto-router kilo-auto/free dynamically picks from the current free pool. โ†ฉ

  4. API-Inference is free for registered users. Current published limits are 2,000 requests/day per user (total across models), with per-model daily quotas dynamically adjusted and capped at 500; concurrency is also dynamically rate-limited. Requires Alibaba Cloud account binding and real-name verification (limits, intro). โ†ฉ

  5. Ollama Cloud measures usage by GPU time, not tokens or requests. Free tier described as "light usage" with session limits resetting every 5 hours and weekly limits every 7 days. Pro (50x more) and Max (250x more) plans available. Not OpenAI SDK-compatible; uses the Ollama API. โ†ฉ

  6. Free models default to 50 RPD per model. A one-time purchase of $10+ in credits unlocks 1,000 RPD for free models. OpenRouter also offers a Free Models Router (openrouter/free) and model fallbacks for chaining models in priority order. Free providers may log prompts for training. โ†ฉ

  7. OVHcloud AI Endpoints offers a permanent free anonymous tier (2 requests per minute per IP, per model) with no signup or API key required. Higher rate limits (400 RPM per Public Cloud project per model) require an API key and are billed pay-as-you-go per token; new Public Cloud accounts get up to $200 in free trial credits. Models are hosted in EU data centers. โ†ฉ

  8. SambaNova grants $5 in initial credits (valid 30 days) on top of the permanent free tier. The free tier itself persists indefinitely with 20 RPM, 20 RPD, and 200K TPD per model. No credit card required. OpenAI SDK-compatible. โ†ฉ