Ollama on a MacBook as an always-on model server
Updated 7 October 2026 · KirsuLab
A MacBook with 16 or more gigabytes of memory is a perfectly good local model server for a household or a small team. What it is not, out of the box, is an always-on one. Left alone it sleeps, and a sleeping server answers nothing. This page is about keeping it up: on a desk with the lid open, on a shelf with the lid closed, and without cooking it.
The setup, in brief
Ollama runs as a menu bar app or as ollama serve in a terminal. Either way it listens on port 11434 and loads models on demand. Three settings matter for a server that other things talk to:
OLLAMA_HOST=0.0.0.0makes it listen on every interface instead of localhost only, so other devices on your network can reach it at the Mac's address.OLLAMA_KEEP_ALIVEsets how long a model stays loaded after the last request. The default is five minutes. A longer value, or-1for forever, trades memory for an instant first reply.OLLAMA_MAX_LOADED_MODELSif more than one model should stay resident at once.
These go into the environment Ollama starts with. For the menu bar app that means launchctl setenv before launch, for a terminal just a prefix on the command, and for a launch agent the EnvironmentVariables block in its plist.
Keep the server on your own network. Nothing in Ollama authenticates requests, and a model server open to the internet will be found and used.
Why it stops: the two sleeps
macOS sleeps a Mac for two unrelated reasons.
Idle sleep. No keyboard or mouse for a while, so the system assumes nobody needs it. A server process does not count as activity. Any app can hold this off with a power assertion, and that is all most people need for a Mac on a desk.
Clamshell sleep. The lid closed. This fires below the assertion system and ignores it. macOS keeps a closed MacBook running only with an external display and the power adapter attached. Otherwise the lid wins, whatever is running.
So a MacBook on a desk needs one thing, and a MacBook on a shelf with the lid shut needs two.
Lid open: hold idle sleep off
Pick one of these.
System Settings. Battery or Energy, then Prevent automatic sleeping on power adapter when the display is off. Simple, global, and it stays that way until you change it back.
caffeinate. Wrap the server in it: caffeinate -i ollama serve. The assertion lives as long as the command does, and dies with the terminal.
A watched process. Valpas has a trigger, While a process is running, that takes names. Type ollama and the Mac stays awake while the server runs, from whichever place it was started, and is free to sleep when you quit it. Choose Allow display to sleep so the screen goes dark and locks on schedule while the Mac keeps serving. The battery guard ends the session if the adapter is pulled and the charge runs low.
Lid closed: get past clamshell sleep
An external display and the adapter. Apple's own clamshell mode. An HDMI dummy plug, a few dollars, stands in for a monitor if there is none. The officially supported route.
pmset. sudo pmset -a disablesleep 1 turns clamshell sleep off entirely, and sudo pmset -a disablesleep 0 turns it back on. It has no timer and survives reboots, so the Mac will not sleep on battery in a bag either until you undo it. The full story.
A session that gives sleep back. The direct download of Valpas has Keep running with the lid closed, free for seven days, then with a Pro key. Start a session, shut the lid, and sleep returns on its own when the session ends, when the adapter is unplugged, when Valpas quits, and when macOS reports serious thermal pressure. Combine it with the process trigger and the hold lasts exactly as long as Ollama runs.
Heat
Inference is heavy work and a closed MacBook breathes through the hinge, not the keyboard. Hard, open surface. Power adapter in. Nothing on top or underneath, never a bed, a cushion or a bag. If the fans run flat out for hours, open the lid an inch or move the machine somewhere cooler. The thermometer in the Valpas panel shows the pressure macOS reports, and Pro releases the lid hold when it reaches serious.
Reaching it
From another device on the same network, the server is at http://<mac address>:11434. Give the Mac a fixed address in your router so it does not move. If you want it from outside the house, use a VPN such as Tailscale or WireGuard rather than opening the port.
When a Mac mini is the right answer
Everything above works. It is also a list of things a desktop does not need. A Mac mini has no lid, no battery, better cooling, and sleeps only when told. If the server becomes a fixture, that is where it ends up. Until then, the MacBook you have is the right place to start.
Related: Keep your Mac awake while an AI agent runs, for the coding agent case.
Questions & answers
Why does my Ollama server stop answering after a while?expand_more
Almost always because the Mac went to sleep. A request to a sleeping Mac times out, and a MacBook sleeps on idle even while a server process is running, because a process is not activity. Keep the Mac awake and the server stays up. Separately, Ollama unloads a model five minutes after its last request by default, so the first reply after a pause is slow, not missing.
Can I run Ollama on a MacBook with the lid closed?expand_more
Yes, but the lid needs its own answer. Power assertions, caffeinate and keep-awake apps do nothing once the lid shuts. You need an external display with the power adapter, or the hidden pmset disablesleep setting, or an app such as Valpas Pro that sets it for a session and puts it back.
How do I reach Ollama from another computer on my network?expand_more
By default it listens on localhost only. Set OLLAMA_HOST to 0.0.0.0 in the environment Ollama starts with, restart it, and point the other device at the Mac's address on port 11434. Keep that to your own network. Exposing a model server to the internet without authentication is a bad idea.
Will a closed MacBook overheat while serving models?expand_more
It can. Inference runs the GPU hard and a shut lid blocks part of the airflow. Hard surface, vents clear, power adapter in, nothing on top. Prefer a tool that gives sleep back when macOS reports serious thermal pressure. A Mac mini has none of this problem, which is why many people end up with one for this job.
Is a Mac mini better than a MacBook for a local model server?expand_more
For an always-on server, yes. No lid, no battery, better cooling, and it sleeps only when you tell it to. A MacBook you already own is the better starting point, and everything on this page is about making that work.
Does Valpas see the ollama process?expand_more
Yes. The While a process is running trigger takes names, so type ollama and the Mac stays awake while the server runs, whether it was started from the menu bar app, a terminal or a launch agent. Pair it with Allow display to sleep so the screen can go dark.