Catalog
49 open-source models, all of them on the phone.
From 112 MB to 2.74 GB, published by 14 labs. Filter by provider, use case or size, then download the ones you want and keep the files.
49 models
All open-source. All run entirely on your device.
Qwen
· 14Qwen 3.5 4B
Latest Qwen with hybrid architecture for detailed answers and multilingual conversations in 201 languages.
2.74 GBNewMultilingualQwen 3.5 2B
FeaturedEfficient hybrid model for everyday tasks with wide language support (201 languages).
1.28 GBNewMultilingualQwen 3.5 0.8B
Compact hybrid model for quick tasks in 201 languages under 600MB.
533 MBNewMultilingualQwen 3 VL 2B
FeaturedAnalyzes images, reads text from screenshots, and answers questions about photos in 32 languages.
1.11 GBVisionThinkingQwen 3 4B
Handles complex questions, multi-step instructions, and conversations in 29+ languages.
2.5 GBMultilingualWeb SearchQwen 3 VL 4B
Advanced image analysis for charts, diagrams, documents, and complex visual reasoning tasks.
2.5 GBVisionThinkingQwen 2.5 Coder 3B
Writes, explains, and refactors code in 92+ programming languages.
2.1 GBCodingQwen 3 4B Thinking
Shows step-by-step reasoning for math, logic puzzles, and complex problems.
2.5 GBThinkingQwen 2.5 3B
Handles translation, creative writing, and longer documents in 29+ languages.
2.1 GBMultilingualWeb SearchQwen 3 1.7B
Balanced model for everyday tasks and multilingual conversations.
1.83 GBMultilingualQwen 2.5 Coder 1.5B
Fast code explanations, bug fixes, and completions in 92+ programming languages.
1.12 GBCodingQwen 2.5 1.5B
Everyday conversations, writing help, and basic tasks in 29+ languages.
1.12 GBMultilingualQwen 3 0.6B
Handles quick questions, simple conversations, and basic tasks.
444 MBQwen 2.5 0.5B
Ultra-fast responses for very simple questions when speed matters most.
491 MBSmall
Meta
· 3Llama 3.2 3B
FeaturedStrong reasoning and detailed explanations for complex conversations in 8 languages.
2.0 GBMultilingualRAM EfficientLlama 3.2 3B Uncensored
Same as Llama 3.2 3B but answers all questions without refusals or content filters.
2.0 GBUncensoredMultilingualLlama 3.2 1B
Fast and lightweight for summarization, Q&A, and basic tasks in 8 languages.
0.8 GBMultilingualRAM Efficient
Microsoft
· 2Phi-4 Mini
Excels at math, logic, and reasoning - competes with much larger models on benchmarks.
2.49 GBNewPhi-3 Mini
Solid math and logic performance trained on high-quality textbook data.
2.2 GB
Gemma 3 4B
FeaturedReasoning and chat model with efficient memory usage for general conversations.
2.49 GBRAM EfficientMultilingualGemma 3 4B QAT
Top-tier reasoning and chat quality with QAT-optimized outputs at same size. Great for general conversations.
2.49 GBNewRAM EfficientGemma 3 1B
Supports 140+ languages in under 1GB - great for less common languages.
806 MBMultilingualGemma 2 2B
Reliable general chat with precise instruction following.
1.71 GBCodeGemma 2B
Specialized for code completion and fill-in-the-middle tasks.
1.63 GBCodingGemma 3 270M
Extremely small (253MB) for very basic tasks when storage is critical.
253 MBSmall
DeepSeek
· 2DeepSeek R1 1.5B
Shows detailed step-by-step reasoning before answering. Great for learning.
1.12 GBThinkingDeepSeek Coder 1.3B
Fast code help in 87 languages. Lightweight but has older training data.
870 MBCoding
Stability AI
· 3Stable Code 3B
Writes and explains code in 18 languages with good context for code review.
1.70 GBCodingStableLM Zephyr 3B
Reliable general chat with consistent, helpful responses.
1.71 GBStableLM 2 Zephyr 1.6B
Conversational model supporting 7 European languages under 1GB.
983 MBMultilingual
Mistral AI
· 1Ministral 3B
FeaturedDetailed answers and strong instruction following in multiple languages.
2.15 GBMultilingualWeb Search
01.AI
· 1Yi Coder 1.5B
Fast code completions and debugging in 52 languages under 1GB.
960 MBCoding
NVIDIA
· 1Nemotron Mini 4B
Strong chat performance for longer conversations.
2.70 GBWeb Search
IBM
· 3Granite 4.0 Micro
FeaturedHandles complex multi-step tasks with strong instruction following and low memory usage.
1.94 GBMultilingualRAM EfficientGranite 4.0 1B
Reliable performance for business and professional tasks under 1GB.
901 MBMultilingualRAM EfficientGranite 4.0 350M
Basic tasks with multilingual support at only 223MB.
223 MBSmall
Hugging Face
· 4SmolLM3 3B
FeaturedPunches above its weight with strong answers and instruction following under 2GB.
1.92 GBNewWeb SearchSmolLM2 1.7B
Good conversations and instruction following at just over 1GB.
1.06 GBSmolLM2 360M
Quick responses and basic reasoning at only 390MB.
390 MBSmallSmolLM2 135M
Tiny model for quick testing. Download a featured model for better responses.
112 MBSmall
OpenAI
· 1GPT-2
Historic GPT-2 from 2019 - the model that started the modern AI wave. Text completion only (not chat). Interesting for educational purposes and AI history.
113 MBSmall
Liquid AI
· 2LFM2.5 1.2B
FeaturedFast and smart in a tiny package. Handles everyday questions and tasks with surprisingly good quality.
731 MBNewRAM EfficientLFM2.5 1.2B Thinking
Works through problems step by step before answering, in a package small enough for any phone.
697 MBThinkingRAM Efficient
Community
· 6Hermes 3 3B
Roleplay and creative writing with consistent character voices.
2.02 GBRoleplayWeb SearchDolphin 2.6 3B
Uncensored model with no content filters for open conversations.
1.79 GBUncensoredRoleplayDanube 3 4B
Balanced 4B model with fast responses for general tasks.
2.39 GBOrca Mini 3B
Shows step-by-step reasoning - trained on GPT-4 thinking traces.
1.98 GBRAM EfficientTinyLlama 1.1B
Basic chat under 700MB - runs fast on almost any device.
670 MBRAM EfficientDanube 3 500M
Ultra-fast responses for basic tasks with multilingual support.
550 MBSmall
Where to start
How to pick one.
Size
18 of the 49 are under a gigabyte.
The smallest is 112 MB and the largest is 2.74 GB. A bigger file generally answers better and generates more slowly, so on an older iPhone the small end is where to start.
Vision
Qwen 3 VL 2B and Qwen 3 VL 4B read images.
They answer questions about a photo or a screenshot, including the text inside one. Everything else in the catalog is text in, text out.
Downloads
One download over Wi-Fi, then it stays.
The file lives on your device from then on, and every answer is generated there, on the phone's CPU and GPU. Nothing is uploaded while you use it.
Start with one model.
Free, no account, and once a model is downloaded it answers with the radios off.
Download KernelAI49 models. 112 MB to 2.74 GB. iOS 15.1 and later.