← Back to iRun

Documentation

User setup instructions for iRun Studio on Mac, iPhone, and iPad. Learn how to load models, browse Language/Vision/Logic models, tailor inference modes, configure tool calling, and manage subscriptions.

User Setup Instructions

Welcome to iRun Studio! Here are quick instructions to get you started:

  • Local LLMs (coding, chatting, debating, thinking, analyzing, vision): download models directly through the built-in Hugging Face browser in the app. If you prefer manual loading, download GGUF or MLX models from huggingface.co and place them in Finder > iRun Docs > LLM-models.
  • Cloud & OpenRouter: connect optional cloud models (Claude, Gemini, DeepSeek, OpenAI, Grok) with zero lock-in.
  • Image models: download CoreML or safetensors models directly or place them in the image models folder.

Quick links (Image Models)

iOS / iPadOS: stable-diffusion-v1-5-GGUF

You can adjust context limits, quantization preferences, and tool settings in the Settings drawer at any time.

Model Browser Architecture

The iRun Studio right sidebar features a categorized Model Browser:

  • Language: General reasoning, coding, conversational LLMs (Gemma, Qwen, DeepSeek, Llama, Bonsai).
  • Vision: Multimodal models capable of image inspection, OCR, and visual question answering.
  • Logic: Specialized mathematical and structured thinking models.

The footer detects all available local weights (e.g. Found 21 GGUF models, Found MLX models) with instant model switching.

Loading & Unloading Models

Step 1: Identify RAM amount (go into iRun > Settings (gear icon) > scroll to see total RAM amount).

Step 2: Downloading & loading. Allocating around 50% of total unified memory is a good guideline for LLMs. Try sticking with Q4_K_M or Q8_K_M quantizations.

  • Q8_0 quants are great for accuracy and preserving deep context.
  • Q4_0 / Q4_K_M quants balance speed and memory efficiency. They retain ~90%+ of unquantized accuracy while cutting RAM usage in half.

Choosing parameter sizes

Generally, a higher parameter model with a Q4 quantization (e.g., 8B Q4) offers greater knowledge depth than a smaller model at full precision (e.g., 2B FP16). Choose the configuration that best fits your unified memory:

  • 8 GB RAM (Mac / iPhone): stick to a 4B–9B model at Q4_0
  • 16 GB RAM: 9B–14B model at Q4_0 or 8B at Q8_0
  • 32 GB+ RAM: 14B–35B models at Q4_0, or 70B models at Q4_K_M on 64GB+ Macs

Newer model architectures (such as Gemma 3, Qwen 3.5, and DeepSeek) deliver remarkable reasoning even at smaller parameter counts.

Choosing the Right AI

Chat & General Reasoning

GGUF format on macOS and iOS. MLX supported on Apple Silicon Macs.

  • Google Gemma: Exceptional all-around performance, fast token throughput, very lightweight footprint.
  • Qwen: Outstanding multi-step reasoning, coding, and mathematical capabilities.
  • Llama: Reliable instruction following and broad world knowledge.
  • Bonsai 8B: Specialized on-device reasoning model by PrismML with zero cloud requirements.

Image Generation

GGUF for iOS/iPadOS; macOS supports both GGUF and safetensors format.

  • Stable Diffusion (SD1.5): Fast on-device diffusion with low memory consumption.
  • SDXL 1.0: High-resolution generation on Apple Silicon M-series.
  • Turbo-SD / Turbo-SDXL: Real-time on-device generation with minimal step count.

Coding, Research & Autonomous Agents

  • Qwen & DeepSeek: High accuracy for code generation, bug hunting, and structured JSON output.
  • Google Gemma: Fast, concise responses ideal for real-time tool execution.

Inference Modes & Tool Calling

iRun Studio allows you to switch inference modes on the fly via bottom pills:

  • ⚡ Fast: Low latency, direct responses with minimal preamble.
  • 🧠 Thinking: Deep chain-of-thought reasoning before output generation.
  • ✨ Pro: Maximum context budget with autonomous tool calling.

Integrated Web Search & Tools: When enabled, iRun automatically executes tools (e.g. web_search) with real-time timers and structured citation cards rendered inline.

Customization

Themes

iRun includes multiple dark and light themes with optional ambient canvas animations. Background particle effects can be paused for maximum battery life.

Manager & Personas (Pro)

Create unlimited custom personalities with unique system prompts, temperature sliders, and tool access profiles.

Universal Features Across All Devices

All power workstation features in iRun Studio are fully available across Mac, iPhone, and iPad with unified Apple Silicon hardware acceleration and Universal Purchase:

  • Autonomous Agent Mode: RAG retrieval, file I/O, code execution, and autonomous tool loops inside a secure sandbox.
  • DeepSearch: Automated multi-source research agent that compiles comprehensive findings into structured PDF reports.
  • CoreML Image Generation: On-device Stable Diffusion & SDXL generation directly on Apple Neural Engine.
  • Revolver Mode: Autonomous loops alternating between plan, build, and bughunt.
  • 120Hz ProMotion Subagent Bar: Real-time task inspection and multi-agent coordination on iPhone, iPad, and MacBook displays.

How can I cancel my subscription?

iRun does not handle billing directly. All subscriptions are securely managed by Apple through your App Store account:

  1. On your Apple device, ensure you are signed into your Apple ID.
  2. Go to device Settings and tap your name card at the top.
  3. Tap Subscriptions (or Media & Purchases > Manage Subscriptions).
  4. Select iRun and tap Cancel Subscription.

Local chat remains 100% free forever even without an active subscription!

Age Restrictions

iRun is intended for developers, researchers, and hobbyists. Because local LLMs generate unmoderated raw outputs directly from weights, guardian supervision is advised for younger users.

App Store Age Rating: 16+