[01/01]
>

And a 35B Model on an iPhone

Run an 80B Qwen in 4.3 GB RAM

Swiftlet Project

02

The Memory Challenge

Running massive LLMs on consumer devices

4.3 GB
RAM needed to run an 80B parameter model
Previously required 160+ GB of memory

Key Technical Achievements

  • 01
    Advanced Quantization Compress model weights with minimal quality loss
  • 02
    Swift-Based Engine Native Apple Silicon optimization for maximum performance
  • 03
    On-Device Processing No cloud dependency — complete privacy and offline capability
  • 04
    Real-Time Generation Responsive inference on consumer-grade hardware
05

Mac & iPhone

Cross-platform AI inference

Device Capabilities

Mac
  • Run 80B parameter models
  • Only 4.3 GB RAM required
  • Full feature support
  • Desktop-grade performance
iPhone Mobile
  • Run 35B parameter models
  • Mobile-optimized inference
  • Portable AI in your pocket
  • On-the-go intelligence

How to Use Swiftlet

1

Clone Repository

Get the code from GitHub

2

Install Framework

Set up Swiftlet on your device

3

Download Weights

Acquire quantized model files

4

Run Inference

Start generating locally

Explore Swiftlet

Bring massive LLMs to your personal devices

github.com/leonickson1/Swiftlet
Made with AirSlide
𝕏 in