Running Local AI with Ollama on Apple Hardware
Getting Started
I've been experimenting with running AI models locally on my Apple hardware, and the experience has been surprisingly smooth and empowering.
My setup is an M4 Pro MacBook Pro with 48GB of unified RAM, which I bought specifically because I wanted to experiment with running LLMs locally. I picked it up in Portland, OR about a year ago after finding a great deal. The key tool for getting started was Ollama — an open-source framework that makes running large language models on your local machine practically effortless.
Here's my journey so far, from my first flight test to running local agentic workflows.
Ollama: The Quick Start
Ollama's primary strength is simplicity. Install it, pull a model, and you're running a chat interface in under a minute:
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Pull a model
ollama pull llama3.1:8b
# Start chatting
ollama run llama3.1:8b
No API keys, no cloud accounts, and no recurring token costs. The model downloads automatically, and you're chatting with a local model offline.
The Models That Worked Well
My journey through local models evolved as newer, more specialized architectures were released:
llama3.1:8b— The first model I ran with Ollama. I downloaded it and spent my flight back home to Monterrey experimenting with it offline. It was a great introduction to local LLMs.devstral:24b— After taking a short break, I resumed experimenting about 3 months ago withdevstral:24b. The step up in reasoning capability and code generation was noticeable.qwen3.6:35b-a3b-coding-nvfp4— About 7 weeks ago, I discovered this model and it has been an absolute game changer. It performs exceptionally well on coding tasks across Python, web development, and cloud infrastructure. I even used it to run some of my first Go code!
On an M4 Pro with 48GB of unified memory, running 24B to 35B quantized models strikes the perfect sweet spot between speed and reasoning quality.
Ollama.cpp & MLX: Deep Tech Stack
Aside from standard Ollama, there are two underlying/alternative runtimes worth keeping an eye on:
- Ollama.cpp / llama.cpp — A lighter-weight, C++-based implementation optimized for direct Metal API calls on Apple Silicon.
- MLX — Apple's native machine learning framework designed specifically to leverage unified memory without CPU/GPU transfer overhead.
Note: I have yet to experiment deeply with Ollama.cpp and MLX natively, but I plan to dive into both soon and will write a dedicated post specifically covering performance benchmarks and setup for both.
Moving to Locally Run Agents
Once I had solid local models running, I started experimenting with local agentic workflows:
- Aider +
devstral:24b: I deliberately avoided OpenClaw and went straight into using Aider paired with Devstral. While the concept was interesting, it felt just okay — I really struggled to build something truly practical or useful with that combination. - Claude Code +
qwen3.6:35b-a3b-coding-nvfp4: Next, I switched to using Claude Code backed by Qwen 3.6 35B. While it wasn't blazing fast, the execution quality and context handling were much, much better and significantly more reliable for actual dev tasks.
What's Next
Local AI is evolving rapidly. I'm keeping a close eye on newer open models trending on Hugging Face — particularly models like Kimi-k3 or DeepSeek v4 Flash — to see if they deliver better speed and accuracy on Apple hardware.
More updates to come as I benchmark new models and test custom agent setups!