Developer · Local MCP server
Llm Kosh
Local-first AI memory cartridge with persistent MCP memory for Claude via SQLite FTS5.
What the MCP Registry states
The entry as published to the official MCP Registry (read 2026-10-04), latest version.
- Registry name
io.github.rastogivaibhav/llm-kosh- Version
- 2.1.3
- Status
- Active
- Category
- developer
- Transport
- stdio (local process)
- Package
- PyPI
- Published
- 2026-06-23
- Updated
- 2026-06-23
- Publisher
- rastogivaibhav (GitHub)
- Repository
- github.com/rastogivaibhav/llm-kosh
- Source
- Registry API entry
Packages
| Registry | Package | Version | Transport |
|---|---|---|---|
| PyPI | llm-kosh | 2.1.3 | stdio |
How to connect Llm Kosh
Llm Kosh runs locally from a Python package published to PyPI: llm-kosh version 2.1.3. It speaks MCP over stdio, so the client starts it as a program and talks to it through standard input and output. It needs Python; clients usually start it with uvx (from uv) or after pip install — the usual command is uvx llm-kosh. It reads the environment variable CARTRIDGE_WORKSPACE (required); set it in the client's configuration for this server.
In the mcpServers JSON format that many desktop and editor MCP clients read, the entry looks like this (placeholders in angle brackets):
{
"mcpServers": {
"llm-kosh": {
"command": "uvx",
"args": [
"llm-kosh"
],
"env": {
"CARTRIDGE_WORKSPACE": "<value>"
}
}
}
}Derived from the registry entry, not tested here. What the server does, and on what terms, is set by its publisher; check its repository or website before giving it access to your accounts or files. How to add an MCP server to an assistant · Before you connect
More developer servers
| Server | Runs |
|---|---|
| LLM CLI GatewayOne MCP endpoint for Claude Code, Codex, Gemini, Grok and Mistral CLIs, with durable async jobs. | Local · stdio |
| LLM ConfiguratorRead-only: which local LLMs a GPU or Mac can run - VRAM fit, tokens/sec, model specs. | Remote · HTTP |
| Llm Cost EstimatorToken counting & multi-model LLM cost estimates: GPT-4o, Claude, Gemini, 25+. No API key. | Local · stdio |
| LLM Hosting PricingLLM and GPU rental prices: model price lookup, GPU listings, cheapest-GPU search, price history. | Remote · HTTP |
| LLM Latency TrackerMeasured latency, time to first token and uptime for ~45 AI inference APIs, by region. | Remote · HTTP |
| Llm Observability OrchestrationRun a prompt through a LangChain (system + human) chain over Gemini on Vertex AI; optional LangSmith. | Remote · HTTP |
| Llm Orchestration AgentRun a prompt through a LangChain (system + human) chain over Gemini on Vertex AI; optional LangSmith. | Remote · HTTP |
| LLM Provider MCPDelegate asynchronous coding jobs between Claude Code, Codex, Cursor Agent, and Pi. | Local · stdio |