Trunk

Trunk

by TuskerLabs

Local inference engine for LLMs on your phone.

Search Hugging Face, download a GGUF model straight to your device, and chat with it completely offline. No cloud, no server, and no account required. Nothing you type is ever sent anywhere.

Android

One app, four moving parts

Models

Search Hugging Face or browse a live Popular list, filtered by default to models sized for your device's memory. Download a GGUF model straight to your phone, or import one you already have. A compatibility badge is computed live from your device's own memory.

Projects

Bind a downloaded model to your own generation settings: temperature, top-p, top-k, context length, and max tokens. Review per-session token and speed history any time.

Playground

Build reusable Agents, a name plus a system prompt, hand-written or drafted by the loaded model from a plain-language description. Wire several into a linear Flow and watch each step run live.

This is our USP →

Inference

Chat with a project's model, or run a saved Flow instead. Multiple auto-titled sessions per project (rename or delete any of them), full markdown with offline LaTeX math and Mermaid flowchart rendering, optional Hexagon NPU or Adreno GPU acceleration on supported Qualcomm chipsets, all running entirely on-device.

Chain agents into a Flow, not just one prompt

Wire reusable Agents together on a touch canvas: tap an output handle, then an input handle, to connect two nodes in order. Run the whole chain on-device and watch each step complete before the final result.

Explainer
Formatter

Nothing leaves your phone.

Every model runs through llama.cpp compiled directly into the app. There's no backend in the loop for inference or model storage, so there's nothing to intercept.

0
Servers your chats touch
0
Accounts required
100%
On-device inference

This describes Trunk today. On-device is the default and always will be; if optional external model providers are ever added, they'll be opt-in, never a requirement.

From install to first chat

1
First launch
A short intro and one-time setup for theme, palette, and stats display.
2
Home
Your device's RAM and storage, plus a suggested max model size.
3
Models
Search, download, or import a GGUF model.
4
Projects
Bind a model to your own generation settings.
5
Playground
Chain reusable Agents into a Flow.
6
Inference
Chat with a project's model, or run a Flow, fully offline.
12:34
Trunk
by TuskerLabs
Welcome to Trunk
Run large language models fully offline, entirely on your own device.
Skip Next
Home
i SETTINGS
This Device
10.9 GB RAM
Suggested model size
up to ~8.3B parameters (Q4)
Library
3 models downloaded
Models
GUIDE SETTINGS
+ Add a model
llama-3.2-1b-instruct-q8_0.gguf
1.3 GB · Downloaded
qwen2.5-coder-3b-q4.gguf
2.1 GB · Downloaded
Projects
GUIDE SETTINGS
+ New Project
Test
llama-3.2-1b · temp 0.7
Coding Helper
qwen2.5-coder-3b · temp 0.3
Playground
GUIDE SETTINGS
+ New Flow or Agent
Write & Review
Summarizer Reviewer
Inference
GUIDE SETTINGS
Project
Coding Helper ›
Chat
New Chat ›
Model
qwen2.5-coder-3b ›
Flow
Direct chat ›
Explain quicksort briefly
Quicksort picks a pivot, partitions the array around it, then recurses on each half.
Ask anything...
12:34
Home
i SETTINGS
This Device
10.9 GB RAM
Suggested model size
up to ~8.3B parameters (Q4)
Library
3 models downloaded

Same design system. Live.

Trunk ships six accent palettes plus Greyscale, each independently tuned for light and dark mode, and a custom hex picker if none of them fit. This entire site runs on the same tokens. Tap a swatch below, it's the exact palette system used inside the app.