LM Studio is a desktop client for local LLMs: download models, manage versions and chat, all in one graphical interface. It can also run in local server mode, giving your own code an endpoint on the same machine.
Where It Fits in the Local AI Toolchain
Running local models usually splits into three parts: the engine (Ollama, llama.cpp, vLLM), the interface (Open WebUI, LM Studio) and the application (editor plugins, scripts). What sets LM Studio apart is that it bundles engine and interface into one app that works out of the box — no need to understand model formats, quantization levels or server flags first.
That makes it best for getting started and trying models: if you want to know how a 7B or 14B model feels on your hardware, it's the fastest route.
Getting Started in Three Steps
- Install: download the installer for your system from the official site; macOS, Windows and Linux are all supported.
- Download a model: search for models and download them inside the app. Inference runs on llama.cpp, and on Apple Silicon Macs it can also use Apple's MLX; you can switch and update runtimes in the app (⌘ Shift R on a Mac).
- Start chatting: load a model and use it in the chat interface. For a first try, pick a smaller quantized version, check speed and memory use, then move up.
For Developers: A Local API
Turn on the server switch in the app's Developer tab, or run lms server start in a terminal, and LM Studio serves an API on your machine at http://localhost:1234/v1 by default. According to the official docs, it offers several kinds of endpoints:
- OpenAI-compatible endpoints: existing OpenAI clients work once you point the base URL at your machine; it also implements
/v1/responses, so tools like Codex can use a local model. - Anthropic-compatible endpoints, plus LM Studio's own REST API.
- Official SDKs:
lmstudio-pythonandlmstudio-js.
Local servers generally don't check the API key; if a client insists on one, any placeholder value works. If you get "connection refused," the server usually isn't running.
It can also act as an MCP host (added in 0.3.17), connecting MCP servers to local models. Some MCP servers can execute code or read local files, so only install ones from sources you trust.
Running Without the Interface
Version 0.4.0 in January 2026 added llmster, a daemon that runs without a graphical interface and can be deployed on servers or cloud instances. The same release added parallel request handling, with four predictions at a time by default. That means switching to Ollama is no longer the only way to keep a model running as a service once you've chosen it — small teams can run internal services on LM Studio directly.
License: Free for Work, but Not Open Source
Since July 8, 2025, using LM Studio at work no longer requires a separate commercial license — the official blog says teams can simply use the app without filling out a form or contacting the company. Before that, it was free only for personal use. A paid Enterprise tier adds features such as single sign-on.
It is still closed-source software. Its app terms of service license it for personal and internal business purposes only, prohibit redistribution, and don't allow offering it to third parties as a hosted service. If your compliance requirements include an auditable toolchain, or you plan to wrap it into an external service, confirm this first.
Practical Limits
Local inference is capped by hardware: GPU memory and RAM decide how large a model you can load and how long a context you can run. Don't carry expectations from cloud flagship models over to a local 14B — they solve different problems. Model files often run from several to tens of gigabytes, so download times depend on your connection.
When to Switch Tools
- A shared chat interface for several people: add Open WebUI.
- A fully open-source, auditable toolchain: move to Ollama or use llama.cpp directly.
- Production-grade throughput: use an inference engine such as vLLM.
