Govern Local LLM Endpoints
Local model servers like Ollama and LM Studio expose OpenAI-compatible endpoints on developer machines. Cup’n’String can discover these endpoints, control which agents may connect, and decide whether they may be exposed beyond the workstation.
Why this matters
Local LLM endpoints are powerful and easy to run, but they often sit outside standard controls. An OpenAI-compatible endpoint on localhost can be reached by any local tool.
Local endpoints (e.g. on port 11434) are typically unmanaged
OpenAI-compatible local APIs can bypass cloud-centric controls
Sensitive prompts may leave controlled flows
There is often no inventory, identity, policy, or audit for local models
Endpoints can be unintentionally exposed to the network
The Cup’n’String approach
Cup’n’String discovers local inference servers and governs access to them through the compatibility adapter and outbound policy.
Auto-discovers local model servers via loopback ports and process signals
Controls which agents and identities may connect to a local endpoint
Decides, by policy, whether an endpoint may be exposed through a controlled tunnel
Attributes and audits access to local model endpoints
Applies outbound policy to provider traffic where routed through the proxy
How it works
- Step 1Local model server (Ollama / LM Studio)
- Step 2Cup’n’String discovery & classification
- Step 3Tenant policy engine
- Step 4Allowed agents & controlled exposure
- Step 5Audit & policy evidence
Checklist
- Can you discover local model endpoints automatically?
- Can you control which agents may connect?
- Can you prevent unintended network exposure?
- Can you audit access to local models?
- Can you apply role-based policies?
- Can you self-host?
Frequently asked questions
Does Cup’n’String work with local model servers?
Yes. It can discover and govern local model endpoints such as Ollama, LM Studio, llama.cpp-compatible servers, and other OpenAI-compatible local gateways, and control which agents may connect and whether endpoints may be exposed remotely.
Can it detect a local endpoint automatically?
Local inference servers are auto-discovered through loopback ports and process signals, then surfaced for governance.
Does this replace Ollama or LM Studio?
No. Cup’n’String governs access to these endpoints; it does not replace them. Developers keep their local models while security gains inventory, policy, and audit.
Can activity be audited?
Yes. Access to local model endpoints can be attributed and recorded for review.
Bring local model endpoints under governance
Inventory, control, and audit Ollama, LM Studio, and other local LLM endpoints across your fleet.