Skip to main content
Supermemory local needs one model provider to power summaries, contextual chunking, and memory extraction. Pick the tab for whichever one you already have a key for. For the full variable reference (fast/text model overrides, offline setup, tuning), see Configuration.
Fully offline — no API key leaves your machine. Any OpenAI-compatible local runner works the same way (LM Studio, vLLM, llama.cpp server); this is the Ollama version.
.env
gpt-oss:20b is a good default for a laptop-class GPU. Bigger models work if you have the VRAM — set OPENAI_MODEL to whatever you’ve pulled.