← Back to iRun

Documentation

User setup instructions for iRun on Mac, iPhone, and iPad. Learn how to load models, pick the right AI, tailor responses, and unlock Pro features.

User Setup Instructions

Welcome to iRun! Here are quick instructions to get you going:

  • LLMs (coding, chatting, debating, thinking, analyzing, vision, etc.): download models through the HuggingFace interface provided in the app. If you must do it manually, go to huggingface.co, download a text-based model, then place it in Finder > iRun Docs > LLM-models.
  • Image models: currently users must download them manually from HuggingFace.

Quick links (Image Models)

iOS / iPadOS: stable-diffusion-v1-5-GGUF

Generally you should tailor your experience through Settings, although most models will function properly regardless of the settings.

Loading & Unloading Models

Step 1: Identify RAM amount (go into iRun > Settings (gear icon) > scroll to see total RAM amount).

Step 2: Downloading & loading. 50% of RAM amount is a good size for LLMs. Try sticking with Q4_K_M or Q8_K_M quantizations.

  • Q8_0 quants are good for accuracy and preserving contexts.
  • Q4_0 quants balance speed and accuracy. They retain about 80–90% of accuracy while being much smaller.

Choosing the right amount of parameters

350M? 1B? 35 billion? Generally, go for a higher parameter count with a lower quant (Q4_0 for example), because it will have more depth and knowledge than smaller, more precise models. Pick the one that fits and runs nicely within your device's memory.

General recommendations

  • 8 GB RAM: stick to a 4B–9B model at Q4_0
  • 16 GB RAM: stick to a 9B–14B model at Q4_0
  • 32 GB RAM: stick to a 9B–14B model at Q8_0, or 14B–35B models at Q4_0

Newer models usually beat older models (even those with higher quants and parameters), so use the latest models for the best experience.

Choosing the Right AI

Chat / General purposes

GGUF format only. MLX is experimental.

  • Google Gemma models work really well (Gemma2, Gemma3, Gemma4)
  • Qwen works well in most cases (Qwen3, Qwen3.5, Qwen3.6)
  • Llama is arguably the simplest. It does as asked but is not very creative or chatty (Llama2, Llama2.1, Llama3)
  • Bonsai 8B by PrismML is also supported (even on iOS), though this model tends to think a lot and generally takes longer to respond. It struggles with intensive tasks since it is not trained for tool usage or tool calling.

Image Gen

GGUF for iOS/iPad. macOS also supports safetensors format.

  • Stable Diffusion (SD1.5) is great, but not all models are compatible. Many models are ComfyUI-specific conversions, so read carefully before installing.
  • SDXL is also compatible (SDXL 1.0)
  • Turbo-SD / Turbo-SDXL are supported. For iOS/iPadOS this is the best option for quality and speed.

Coding / Research / Heavy Tasks

  • Google Gemma (Gemma3, Gemma4) has proven most capable so far, even at low parameter and quant count. Great for 8 GB RAM devices. Strong overall skills, though they struggle with specific, detailed tasks.
  • Qwen offers more agentic work. Qwen3.6 models are great at analysis, MOE, tailored responses, and code completion/bug fixing. Good for specific, detailed tasks.
  • DeepSeek offers great accuracy, but at the cost of speed.

Response Tailoring

You can tailor how the AI responds in three distinct ways:

  • Fast / Thinking / Pro (free): choose the mode depending on task complexity for more thinking or tool usage.
  • Response sliders: a rough budget the AI allocates to each section (thinking, analysis, chat, code, vision, etc.). It tells the AI what to emphasize more or less.
  • Manager (Pro): custom personalities with custom instructions and per-model profiles. You can have unlimited personalities for the same AI and load only the one you need.

Customization

Themes

The app has 4 themes (including default). Some have animated backgrounds/objects, which can be paused for optimization.

Manager (Pro)

Create and manage custom personalities and agents.

macOS-Only Features

We tried to make the app as cross-platform compatible as possible, but heavy features are reserved for macOS:

  • Tool Usage: Agent receives access to tools
  • Agent Mode (Beta): comprehensive RAG tools, lightweight system prompts, Agentic Harness, Revolver Mode (AI alternates between plan, build, and bughunt automatically)
  • DeepSearch (Testing / Early Access): research tab that finds sources and compiles reports (PDF format)
  • Workflows (Coming Soon): create custom workflows using agents/personalities. More information will be available when the feature nears completion.

How can I cancel my subscription?

iRun does not handle the subscription process. Apple does. Here is what you need to do:

  1. On your Apple device, ensure you are signed into your Apple ID with the account holding the membership.
  2. Go to your device's Settings and tap your username/email card.
  3. In the account menu, look for Media & Purchases (App Store icon). Tap Manage Subscriptions.
  4. Follow Apple's steps to cancel your subscription.

You're welcome back anytime!

Age Restrictions

Due to how LLMs behave and operate, we can't be certain they will always remain safe or child-friendly. This app is made for hobbyists and should not be used by children, at least without the supervision of a guardian.

App Store Age Rating: 16+