Run the Bonsai 2 27B LLM Locally on macOS
Job to be done: Set up and run a compressed LLM (Bonsai 2 27B) locally on macOS
🇳🇬 Ways to use this in Nigeria
Ideas to get you started, adapt to your situation.
- Student
A computer science student can set up this local LLM to experiment with text generation for their final year project, like building a simple chatbot or a content summarizer, without needing internet access or paying for API calls.
- Entrepreneur
A tech founder can run this LLM locally to rapidly prototype AI features for their startup's product, such as an automated customer support response generator or a marketing copy assistant, without cloud API expenses.
- 9-5 employee
A software developer in an organization can use this local LLM to prototype internal tools for text summarization of reports or automated email drafting, ensuring company data privacy and reducing reliance on external AI services.
What you’ll get
You will set up and run the Bonsai 2 27B large language model (LLM) directly on your Apple Silicon Mac. This model is known for its near-lossless compression, meaning it’s much smaller (about 6 GB) than typical LLMs of its size while still performing well. You can interact with it through a simple web interface in your browser or via an API (Application Programming Interface), which allows other programs to talk to it.
Tools you need
- curl (free): A command-line tool used to download files from the internet.
- tar (free): A command-line tool built into macOS for extracting compressed archive files.
- Prism’s llama.cpp (free): A special version of
llama.cpp, an open-source project that allows large language models to run efficiently on your computer’s hardware. This specific version is needed for the Bonsai 2 27B model. - Hugging Face (freemium): A platform that hosts many AI models. You will download the Bonsai 2 27B model file from here.
- uv (free): A fast Python package installer and manager. You will use it to install and run the
llmcommand-line tool. - llm (free): A command-line tool by Simon Willison for interacting with large language models, including those running locally.
Steps
This workflow is designed for Apple Silicon Macs (M-series chips) only. You will use your computer’s Terminal application for all steps.
-
Open Terminal: Open the Terminal application on your Mac. You can find it in
Applications > Utilities > Terminalor by searching for “Terminal” in Spotlight (Command + Space). -
Prepare your workspace: Change your current directory to
/tmp. This is a temporary folder where downloaded files can be stored without cluttering your main directories.# macOS or Linux cd /tmpYou should see your terminal prompt change to reflect the new directory, for example,
/tmp %. -
Download the Prism llama.cpp runtime: Use
curlto download the specific runtime needed for this model. This file is a compressed archive containing thellama-serverprogram.# macOS or Linux curl -fL https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gzYou should see
curldownload the file, showing progress, and then the filebonsai-runtime.tar.gzwill be saved in your/tmpdirectory. -
Extract the runtime: Use
tarto uncompress and extract the downloaded runtime file.# macOS or Linux tar -xzf bonsai-runtime.tar.gzYou should see a new folder appear in
/tmp(e.g.,llama-prism-b10685-7dffb15) containing thellama-serverprogram. -
Download the Bonsai 2 27B model: Download the actual language model file, which is in GGUF format. GGUF is a file format optimized for running LLMs efficiently on consumer hardware.
# macOS or Linux curl -fL https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.ggufThis is a large file (around 6 GB), so it will take some time to download depending on your internet speed. You should see
curlshowing download progress, and then the fileTernary-Bonsai-2-27B-PTQ1_0.ggufwill be saved in/tmp. -
Start the LLM server: Run the
llama-serverprogram to load the model and make it available. The-mflag specifies the model file,--portsets the network port,-ngl 99attempts to offload 99 layers to the GPU for faster processing,-fa onenables flash attention, and-c 32768sets the context window size.# macOS or Linux ./llama-prism-b10685-7dffb15/llama-server \ -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \ --port 8331 -ngl 99 -fa on -c 32768The server will start, and you should see many lines of text in your Terminal as it loads the model. Look for a message indicating the server is listening on
http://localhost:8331. -
Access the web interface: Open your web browser (like Chrome, Safari, or Firefox) and go to the address
http://localhost:8331.localhostrefers to your own computer, and8331is the specific network port the server is using.You should see a web interface for the
llama-server, allowing you to chat with the Bonsai 2 27B model directly in your browser. -
(Optional) Interact via API using the LLM CLI: If you prefer to interact with the model using a command-line tool, you can use
uvandllm. First, ensureuvis installed (e.g., by following instructions onastral.sh/uv). Then, installllmand itsopenaiplugin:# macOS or Linux uv pip install llm uv run llm install openaiOnce installed, you can send a prompt to your local server. The author used
uvx(a custom script for runninguvcommands) for this, but you can achieve the same withuv run:# macOS or Linux uv run llm openai endpoint http://127.0.0.1:8331/v1 \ --model bonsai-2-27b --responses hiYou should see a response from the model printed in your Terminal. The author noted varying speeds, sometimes around 20 tokens per second, sometimes 44 tokens per second.
Original source
This workflow is based on a comment by Simon Willison on Hacker News, where he shared instructions for running the Bonsai 2 27B model using a specific fork of llama.cpp and his llm command-line tool.
Notes & variations
- Hardware requirement: This workflow is specifically designed for Apple Silicon Macs (M-series chips). The downloaded runtime binary (
macos-arm64) will not work on Intel Macs or Windows/Linux computers without significant modifications (e.g., compilingllama.cppfor your specific system or using Windows Subsystem for Linux (WSL) for Linux binaries). - Initial data usage: The model file is approximately 6 GB. Ensure you have a stable internet connection and sufficient data if you are on a metered plan before downloading.
- GPU acceleration issues: The author noted a message “ggml_metal_device_init: - the tensor API is not supported in this environment - disabling” during server startup. This indicates that the model might not be fully utilizing your Mac’s GPU (Graphics Processing Unit) for acceleration, even with the
-ngl 99flag. If you encounter this, the model will still run on your CPU (Central Processing Unit), but it might be slower. There isn’t a simple fix provided in the source, but it’s a common observation withllama.cppforks and specific hardware/software configurations. - Getting better results: If you experience slow performance or the GPU error, you can try adjusting the
-nglparameter (Number of GPU Layers). Reducing it (e.g.,-ngl 0to run entirely on CPU) or removing it might sometimes resolve issues if the GPU offloading isn’t working correctly for your specific system setup. Experimenting with this value can help find the best balance for your Mac.