KernelAI
FeaturesModelsDownloadContact
Get the app
KernelAI
FeaturesModelsDownloadContactGet the app
Home

Catalog

49 open-source models, all of them on the phone.

From 112 MB to 2.74 GB, published by 14 labs. Filter by provider, use case or size, then download the ones you want and keep the files.

Get KernelAI

Provider

Use case

Size

49 models

All open-source. All run entirely on your device.

Qwen

· 14
  • Qwen 3.5 4B

    Latest Qwen with hybrid architecture for detailed answers and multilingual conversations in 201 languages.

    2.74 GB
    NewMultilingual
  • Qwen 3.5 2B

    Featured

    Efficient hybrid model for everyday tasks with wide language support (201 languages).

    1.28 GB
    NewMultilingual
  • Qwen 3.5 0.8B

    Compact hybrid model for quick tasks in 201 languages under 600MB.

    533 MB
    NewMultilingual
  • Qwen 3 VL 2B

    Featured

    Analyzes images, reads text from screenshots, and answers questions about photos in 32 languages.

    1.11 GB
    VisionThinking
  • Qwen 3 4B

    Handles complex questions, multi-step instructions, and conversations in 29+ languages.

    2.5 GB
    MultilingualWeb Search
  • Qwen 3 VL 4B

    Advanced image analysis for charts, diagrams, documents, and complex visual reasoning tasks.

    2.5 GB
    VisionThinking
  • Qwen 2.5 Coder 3B

    Writes, explains, and refactors code in 92+ programming languages.

    2.1 GB
    Coding
  • Qwen 3 4B Thinking

    Shows step-by-step reasoning for math, logic puzzles, and complex problems.

    2.5 GB
    Thinking
  • Qwen 2.5 3B

    Handles translation, creative writing, and longer documents in 29+ languages.

    2.1 GB
    MultilingualWeb Search
  • Qwen 3 1.7B

    Balanced model for everyday tasks and multilingual conversations.

    1.83 GB
    Multilingual
  • Qwen 2.5 Coder 1.5B

    Fast code explanations, bug fixes, and completions in 92+ programming languages.

    1.12 GB
    Coding
  • Qwen 2.5 1.5B

    Everyday conversations, writing help, and basic tasks in 29+ languages.

    1.12 GB
    Multilingual
  • Qwen 3 0.6B

    Handles quick questions, simple conversations, and basic tasks.

    444 MB
  • Qwen 2.5 0.5B

    Ultra-fast responses for very simple questions when speed matters most.

    491 MB
    Small

Meta

· 3
  • Llama 3.2 3B

    Featured

    Strong reasoning and detailed explanations for complex conversations in 8 languages.

    2.0 GB
    MultilingualRAM Efficient
  • Llama 3.2 3B Uncensored

    Same as Llama 3.2 3B but answers all questions without refusals or content filters.

    2.0 GB
    UncensoredMultilingual
  • Llama 3.2 1B

    Fast and lightweight for summarization, Q&A, and basic tasks in 8 languages.

    0.8 GB
    MultilingualRAM Efficient

Microsoft

· 2
  • Phi-4 Mini

    Excels at math, logic, and reasoning - competes with much larger models on benchmarks.

    2.49 GB
    New
  • Phi-3 Mini

    Solid math and logic performance trained on high-quality textbook data.

    2.2 GB

Google

· 6
  • Gemma 3 4B

    Featured

    Reasoning and chat model with efficient memory usage for general conversations.

    2.49 GB
    RAM EfficientMultilingual
  • Gemma 3 4B QAT

    Top-tier reasoning and chat quality with QAT-optimized outputs at same size. Great for general conversations.

    2.49 GB
    NewRAM Efficient
  • Gemma 3 1B

    Supports 140+ languages in under 1GB - great for less common languages.

    806 MB
    Multilingual
  • Gemma 2 2B

    Reliable general chat with precise instruction following.

    1.71 GB
  • CodeGemma 2B

    Specialized for code completion and fill-in-the-middle tasks.

    1.63 GB
    Coding
  • Gemma 3 270M

    Extremely small (253MB) for very basic tasks when storage is critical.

    253 MB
    Small

DeepSeek

· 2
  • DeepSeek R1 1.5B

    Shows detailed step-by-step reasoning before answering. Great for learning.

    1.12 GB
    Thinking
  • DeepSeek Coder 1.3B

    Fast code help in 87 languages. Lightweight but has older training data.

    870 MB
    Coding

Stability AI

· 3
  • Stable Code 3B

    Writes and explains code in 18 languages with good context for code review.

    1.70 GB
    Coding
  • StableLM Zephyr 3B

    Reliable general chat with consistent, helpful responses.

    1.71 GB
  • StableLM 2 Zephyr 1.6B

    Conversational model supporting 7 European languages under 1GB.

    983 MB
    Multilingual

Mistral AI

· 1
  • Ministral 3B

    Featured

    Detailed answers and strong instruction following in multiple languages.

    2.15 GB
    MultilingualWeb Search

01.AI

· 1
  • Yi Coder 1.5B

    Fast code completions and debugging in 52 languages under 1GB.

    960 MB
    Coding

NVIDIA

· 1
  • Nemotron Mini 4B

    Strong chat performance for longer conversations.

    2.70 GB
    Web Search

IBM

· 3
  • Granite 4.0 Micro

    Featured

    Handles complex multi-step tasks with strong instruction following and low memory usage.

    1.94 GB
    MultilingualRAM Efficient
  • Granite 4.0 1B

    Reliable performance for business and professional tasks under 1GB.

    901 MB
    MultilingualRAM Efficient
  • Granite 4.0 350M

    Basic tasks with multilingual support at only 223MB.

    223 MB
    Small

Hugging Face

· 4
  • SmolLM3 3B

    Featured

    Punches above its weight with strong answers and instruction following under 2GB.

    1.92 GB
    NewWeb Search
  • SmolLM2 1.7B

    Good conversations and instruction following at just over 1GB.

    1.06 GB
  • SmolLM2 360M

    Quick responses and basic reasoning at only 390MB.

    390 MB
    Small
  • SmolLM2 135M

    Tiny model for quick testing. Download a featured model for better responses.

    112 MB
    Small

OpenAI

· 1
  • GPT-2

    Historic GPT-2 from 2019 - the model that started the modern AI wave. Text completion only (not chat). Interesting for educational purposes and AI history.

    113 MB
    Small

Liquid AI

· 2
  • LFM2.5 1.2B

    Featured

    Fast and smart in a tiny package. Handles everyday questions and tasks with surprisingly good quality.

    731 MB
    NewRAM Efficient
  • LFM2.5 1.2B Thinking

    Works through problems step by step before answering, in a package small enough for any phone.

    697 MB
    ThinkingRAM Efficient

Community

· 6
  • Hermes 3 3B

    Roleplay and creative writing with consistent character voices.

    2.02 GB
    RoleplayWeb Search
  • Dolphin 2.6 3B

    Uncensored model with no content filters for open conversations.

    1.79 GB
    UncensoredRoleplay
  • Danube 3 4B

    Balanced 4B model with fast responses for general tasks.

    2.39 GB
  • Orca Mini 3B

    Shows step-by-step reasoning - trained on GPT-4 thinking traces.

    1.98 GB
    RAM Efficient
  • TinyLlama 1.1B

    Basic chat under 700MB - runs fast on almost any device.

    670 MB
    RAM Efficient
  • Danube 3 500M

    Ultra-fast responses for basic tasks with multilingual support.

    550 MB
    Small

Where to start

How to pick one.

  • Size

    18 of the 49 are under a gigabyte.

    The smallest is 112 MB and the largest is 2.74 GB. A bigger file generally answers better and generates more slowly, so on an older iPhone the small end is where to start.

  • Vision

    Qwen 3 VL 2B and Qwen 3 VL 4B read images.

    They answer questions about a photo or a screenshot, including the text inside one. Everything else in the catalog is text in, text out.

  • Downloads

    One download over Wi-Fi, then it stays.

    The file lives on your device from then on, and every answer is generated there, on the phone's CPU and GPU. Nothing is uploaded while you use it.

Start with one model.

Free, no account, and once a model is downloaded it answers with the radios off.

Download KernelAI

49 models. 112 MB to 2.74 GB. iOS 15.1 and later.

Back to the homepage
KernelAI

Private AI on your device. 49 open-source models, on your device, free.

Product

  • Models
  • Download

Company

  • Contact
  • Privacy
  • Terms

KernelAI

© 2026 KernelAI