ai-lab

Run real AI models directly in your browser — now including text generation. Model weights download once from Hugging Face and inference happens on your device: free, unlimited, private. No API key, no server, no rate limits.

Model

detecting device…
Generative models: ~90 MB (135M) or ~350 MB (0.5B) one-time download, cached afterwards.

Generation settings

0.7
96

Input

Output

Run the model to see results here.
Why run models in the browser? Hosted inference APIs meter usage and the free tiers keep shrinking — GitHub's own Models inference API was retired in 2026. Local inference is the durable alternative: the model runs on your own hardware, so it costs nothing, cannot be rate-limited, and your prompts never leave the page. A 135M instruct model is genuinely useful for short drafts and lists; a 0.5B model trades a heavier download for a bit more coherence. Both are real transformer models with chat templates, not toys.