Skip to content
Gauge

ax-check.com

LeanCTX

leanctx.com

ax-check.com

Checked 19 September 2026

B

79/100

Agents haven’t verified setup

Docs, install and pricing pages all work without login

20 of 23 checks passed: quickstart, CLI/SDK install, API and MCP docs all render statically. Pro/Team pricing lacks a stated price or unit, only "confirmed before purchase."

3 fixes to start with

Concrete changes, backed by what we found.

Onboarding

Resolve the blocker that stopped agent setup.

Agents cannot progress through onboarding without human help. Provide an agent-accessible path through authentication and setup.

Docs

State what each price includes and how overages are billed.

Pro and Team show no price, unit or limit — only 'Confirmed before purchase'.

Docs

Return Markdown when clients request text/markdown.

Homepage returned text/html even when text/markdown was requested; no Markdown representation served.

Recorded coding sessions

Same task, 3 models, no leanctx.com account. Recorded 19 September 2026 on Pi.

3 recorded · 4m 41s

4m 41s · recorded

Session findings

Select a finding to see it in context.

piv0.84.2
modelQwen 3.8 Maxcwd/sandbox/repo
Help me build a simple example using LeanCTX. Tell me how pricing works, and briefly tell me whether this product will be easy for you to manage. Let me know if you get blocked. If this product has no developer workflow you can act on, say so plainly and stop. Stay light: use the hosted product through its SDK or API. Do not start local service stacks or wait for long-running commands; if the quickstart requires either, say so plainly and stop.
Thinking…

I'm checking whether LeanCTX is available in the repository.

$ ls -la /sandbox/repo && grep -ri "leanctx" /sandbox/repo -l | head -20{"content":[{"text":"total 16\ndrwxr-xr-x 3 root root 4096 Sep 18 08:00 .\ndrwxr-xr-x 7 root root 4096 Sep 19 01:59 ..\ndrwxr-xr-x 8 root root 4096 Sep 18 08:00…ctrl+o
{"content":[{"text":"total 16\ndrwxr-xr-x 3 root root 4096 Sep 18 08:00 .\ndrwxr-xr-x 7 root root 4096 Sep 19 01:59 ..\ndrwxr-xr-x 8 root root 4096 Sep 18 08:00 .git\n-rw-r--r-- 1 root root  408 Sep 18 08:00 .gitignore\n","type":"text"}]}
$ pip list 2>/dev/null | grep -i lean; npm ls -g 2>/dev/null | grep -i lean; ls /usr/local/lib/node_modules 2>/dev/null{"content":[{"text":"@anthropic-ai\n@earendil-works\n@openai\ncorepack\nnpm\n","type":"text"}]}
Thinking…

Empty repo. LeanCTX — is this a real product? I don't know it. It might be fictional. Should I check on the internet? There's probably no network access. Let me quickly check.

$ pip download leanctx --no-deps -d /tmp/x 2>&1 | tail -3; npm view leanctx 2>&1 | head -5{"content":[{"text":"Downloading leanctx-0.3.1-py3-none-any.whl (66 kB)\nSaved /tmp/x/leanctx-0.3.1-py3-none-any.whl\nSuccessfully downloaded leanctx\n\nleanctx…ctrl+o
{"content":[{"text":"Downloading leanctx-0.3.1-py3-none-any.whl (66 kB)\nSaved /tmp/x/leanctx-0.3.1-py3-none-any.whl\nSuccessfully downloaded leanctx\n\nleanctx@0.0.0 | MIT | deps: none | versions: 1\nDrop-in prompt compression for production LLM applications. Cut Anthropic / OpenAI / Gemini bills by 40-60%. Placeholder — TypeScript SDK coming in v0.3.\nhttps://github.com/jia-gao/leanctx\n\n","type":"text"}]}
Thinking…

There's a Python package leanctx 0.3.1. Let's install it and check the SDK.

$ pip install leanctx 2>&1 | tail -2; python -c "import leanctx; print(leanctx.__file__); print(dir(leanctx))"{"content":[{"text":"Installing collected packages: leanctx\nSuccessfully installed leanctx-0.3.1\n/opt/freestyle/python/lib/python3.12/site-packages/leanctx/__…ctrl+o
{"content":[{"text":"Installing collected packages: leanctx\nSuccessfully installed leanctx-0.3.1\n/opt/freestyle/python/lib/python3.12/site-packages/leanctx/__init__.py\n['Anthropic', 'AsyncAnthropic', 'AsyncOpenAI', 'CompressionStats', 'Compressor', 'ContentType', 'DedupStrategy', 'Gemini', 'Lingua', 'MethodStatus', 'Middleware', 'ObservabilityConfig', 'OpenAI', 'PurgeErrorsStrategy', 'RepeatTracker', 'Router', 'SelfLLM', 'Strategy', 'Verbatim', '__all__', '__builtins__', '__cached__', '__doc__', '__file__', '__loader__', '__name__', '__package__', '__path__', '__spec__', '__version__', '_content', '_gemini_adapter', 'classifier', 'classify', 'client', 'compressors', 'middleware', 'observability', 'router', 'stats', 'strategies', 'tokens']\n","type":"text"}]}
Thinking…

I need to look into this further. Let me confirm the details.

$ cd /opt/freestyle/python/lib/python3.12/site-packages/leanctx && ls && head -100 __init__.py{"content":[{"text":"__init__.py\n__pycache__\n_content.py\n_gemini_adapter.py\nbench\nclassifier.py\nclient.py\ncompressors\nintegrations\nmiddleware.py\nobser…ctrl+o
{"content":[{"text":"__init__.py\n__pycache__\n_content.py\n_gemini_adapter.py\nbench\nclassifier.py\nclient.py\ncompressors\nintegrations\nmiddleware.py\nobservability\npy.typed\nrouter.py\nstats.py\nstrategies\ntokens.py\n\"\"\"leanctx — drop-in prompt compression for production LLM applications.\n\nv0.0.x is a name-reservation release with the public SDK surface scaffolded\nas a passthrough. The working release (v0.1) swaps the passthrough for the\nreal three-mode compression pipeline (local LLMLingua-2 / self_llm / hybrid).\n\nSee https://github.com/jia-gao/leanctx for progress.\n\"\"\"\n\nfrom leanctx.classifier import RepeatTracker, classify\nfrom leanctx.client import Anthropic, AsyncAnthropic, AsyncOpenAI, Gemini, OpenAI\nfrom leanctx.compressors import Compressor, ContentType, Lingua, SelfLLM, Verbatim\nfrom leanctx.middleware import Middleware\nfrom leanctx.observability import MethodStatus, ObservabilityConfig\nfrom leanctx.router import Router\nfrom leanctx.stats import CompressionStats\nfrom leanctx.strategies import DedupStrategy, PurgeErrorsStrategy, Strategy\n\n__version__ = \"0.3.1\"\n\n__all__ = [\n    \"Anthropic\",\n    \"AsyncAnthropic\",\n    \"AsyncOpenAI\",\n    \"CompressionStats\",\n    \"Compressor\",\n    \"ContentType\",\n    \"DedupStrategy\",\n    \"Gemini\",\n    \"Lingua\",\n    \"MethodStatus\",\n    \"Middleware\",\n    \"ObservabilityConfig\",\n    \"OpenAI\",\n    \"PurgeErrorsStrategy\",\n    \"RepeatTracker\",\n    \"Router\",\n    \"SelfLLM\",\n    \"Strategy\",\n    \"Verbatim\",\n    \"__version__\",\n    \"classify\",\n]\n","type":"text"}]}
Thinking…

Let me understand the SDK. Is this a hosted product, or a local library? The user said "use the hosted product via its SDK or API." LeanCTX seems to be a local library (LLMLingua-2 compression runs locally). Let me check client.py and compressors to understand, and check the pricing. Pricing: check the GitHub repo / PyPI page. Probably there's no pricing info — this is an OSS library, and its "pricing" is about cutting LLM costs. Let me take a quick look at client.py.

$ cd /opt/freestyle/python/lib/python3.12/site-packages/leanctx && wc -l client.py compressors/*.py middleware.py router.py strategies/*.py && head -80 client.py{"content":[{"text":" 648 client.py\n 20 compressors/__init__.py\n 51 compressors/base.py\n 277 compressors/lingua.py\n 325 compressors/selfllm.py\n 49…ctrl+o
{"content":[{"text":"  648 client.py\n   20 compressors/__init__.py\n   51 compressors/base.py\n  277 compressors/lingua.py\n  325 compressors/selfllm.py\n   49 compressors/verbatim.py\n  261 middleware.py\n   40 router.py\n   20 strategies/__init__.py\n   22 strategies/base.py\n   39 strategies/dedup.py\n   47 strategies/purge_errors.py\n 1799 total\n\"\"\"Drop-in LLM SDK wrappers.\n\nInterface-compatible wrappers around the official anthropic / openai /\ngoogle-genai SDKs. Public surface mirrors the upstream client; an\nadditional ``leanctx_config`` kwarg controls compression behavior, and\neach response gains ``usage.leanctx_*`` telemetry attributes.\n\nWhen the ``[otel]`` extra is installed and observability is enabled\nin ``leanctx_config[\"observability\"][\"otel\"]``, every wrapper request\npath emits a ``leanctx.compress`` span.\n\"\"\"\n\nfrom __future__ import annotations\n\nfrom typing import Any\n\nfrom leanctx._gemini_adapter import (\n    contents_to_messages as _gemini_contents_to_messages,\n)\nfrom leanctx._gemini_adapter import (\n    messages_to_contents as _gemini_messages_to_contents,\n)\nfrom leanctx.middleware import Middleware\nfrom leanctx.observability.config import ObservabilityConfig\nfrom leanctx.observability.middleware_hooks import compression_span\nfrom leanctx.observability.stream_owners import (\n    _SpanOwningAsyncIterator,\n    _SpanOwningIterator,\n)\nfrom leanctx.stats import CompressionStats\n\n_ANTHROPIC_INSTALL_HINT = (\n    \"The 'anthropic' package is required for leanctx.Anthropic. \"\n    \"Install with: pip install 'leanctx[anthropic]'\"\n)\n\n_OPENAI_INSTALL_HINT = (\n    \"The 'openai' package is required for leanctx.OpenAI. \"\n    \"Install with: pip install 'leanctx[openai]'\"\n)\n\n_GEMINI_INSTALL_HINT = (\n    \"The 'google-genai' package is required for leanctx.Gemini. \"\n    \"Install with: pip install 'leanctx[gemini]'\"\n)\n\n\ndef _load_anthropic() -> Any:\n    try:\n        import anthropic\n    except ImportError as e:\n        raise ImportError(_ANTHROPIC_INSTALL_HINT) from e\n    return anthropic\n\n\ndef _load_openai() -> Any:\n    try:\n        import openai\n    except ImportError as e:\n        raise ImportError(_OPENAI_INSTALL_HINT) from e\n    return openai\n\n\ndef _load_gemini() -> Any:\n    try:\n        from google import genai\n    except ImportError as e:\n        raise ImportError(_GEMINI_INSTALL_HINT) from e\n    return genai\n\n\ndef _parse_observability(leanctx_config: dict[str, Any] | None) -> ObservabilityConfig:\n    if not leanctx_config:\n        return ObservabilityConfig()\n    return ObservabilityConfig.from_dict(leanctx_config.get(\"observability\"))\n\n\nclass Anthropic:\n    \"\"\"Drop-in replacement for ``anthropic.Anthropic``.\"\"\"\n\n","type":"text"}]}
$ curl -s https://raw.githubusercontent.com/jia-gao/leanctx/main/README.md 2>&1 | head -120{"content":[{"text":"# leanctx\n\n[![PyPI](https://img.shields.io/pypi/v/leanctx)](https://pypi.org/project/leanctx/)\n[![Python](https://img.shields.io/pypi/py…ctrl+o
{"content":[{"text":"# leanctx\n\n[![PyPI](https://img.shields.io/pypi/v/leanctx)](https://pypi.org/project/leanctx/)\n[![Python](https://img.shields.io/pypi/pyversions/leanctx)](https://pypi.org/project/leanctx/)\n[![License](https://img.shields.io/pypi/l/leanctx)](LICENSE)\n\n**Drop-in prompt compression for production LLM applications.**\nCut your input-token bill by 10–40%,without sacrificing accuracy.\n\n```python\n# before\nfrom openai import OpenAI\n\n# after\nfrom leanctx import OpenAI  # same interface, compressed requests\n```\n\nOn the full **[LongBench v2](https://longbench2.github.io/)** set (N=503), layered on top of a compressor that is already running, leanctx removes **an extra 18.7 % of tokens** — rising to **36.7 % on prose-heavy traffic** — at a cost of 1.8 pp of accuracy. Every figure regenerates from [per-item records committed to this repo](#every-figure-here-regenerates-from-committed-data). Open-source models, runs locally, MIT-licensed. Your prompts and user data never leave your infrastructure by default.\n\n[Quickstart](#quickstart-60-seconds) · [What makes it different](#what-makes-it-different) · [Benchmarks](#real-numbers) · [Integrations](#integrations) · [How it works](#how-it-works)\n\n---\n\n## What makes it different\n\n**Loss-tolerance routing** — classify every segment of a prompt by how much\ndistortion it can survive, then compress each class differently.\n\n```mermaid\nflowchart TD\n    A[\"Agent request\"] --> B[\"Classify by loss tolerance\"]\n    B -->|zero tolerance| C[\"Verbatim<br/>byte for byte\"]\n    B -->|high tolerance| D[\"LLMLingua-2<br/>~50% removed\"]\n    B -->|opt-in| E[\"Self-LLM\"]\n    C --> F[\"Recompose + check invariants\"]\n    D --> F\n    E --> F\n    F -->|hold| G[\"Compressed → provider\"]\n    F -->|fail or timeout| H[\"Original → provider\"]\n\n    classDef zero fill:#e8f1fd,stroke:#2a78d6,color:#0b0b0b\n    classDef high fill:#e9f4f0,stroke:#1baf7a,color:#0b0b0b\n    classDef open fill:#f4f4f1,stroke:#898781,color:#0b0b0b\n    class C zero\n    class D high\n    class H open\n```\n\n| Class | Content | Treatment | Effect |\n|---|---|---|---|\n| **Zero tolerance** | code, stack traces, `tool_use_id`, tool name and input, JSON | verbatim | 0 % altered |\n| **High tolerance** | documentation, retrieved passages, logs, prior turns | LLMLingua-2, on-device | ~50 % removed |\n| **Conditional** | low-confidence prose, oversized context | self-LLM, opt-in | 41–49 % removed |\n\nCompression applied uniformly to a prompt will eventually damage the one part that\ncannot survive being touched, and will do so unpredictably. The hard part is not\nshortening text — it is deciding, inside a live request, what may be shortened at all.\n\n---\n\n## Quickstart (60 seconds)\n\n```bash\npip install 'leanctx[openai,lingua]'    # or [anthropic], [gemini]\n```\n\n```python\nfrom leanctx import OpenAI\n\nclient = OpenAI(\n    leanctx_config={\n        \"mode\": \"on\",\n        \"trigger\": {\"threshold_tokens\": 2000},\n        \"routing\": {\"prose\": \"lingua\"},  # route prose through LLMLingua-2\n    },\n)\n\nresponse = client.chat.completions.create(\n    model=\"gpt-4o-mini\",\n    max_tokens=512,\n    messages=[{\"role\": \"user\", \"content\": LONG_DOCUMENT}],\n)\n\nprint(response.usage.leanctx_tokens_saved)  # e.g. 1841\nprint(response.usage.leanctx_ratio)         # e.g. 0.49\n```\n\n> **First Lingua call loads ~1.2 GB of model weights** to `~/.cache/huggingface/`. Subsequent calls reuse the cache. Add `pip install 'leanctx[lingua]'` to opt in; without it, leanctx falls back to passthrough.\n\nVerify the install with no API key needed:\n\n```bash\nleanctx bench list                                   # 7 registered scenarios\nleanctx bench run agent-structural --workload agent  # 5 invariants enforced, exit 0 = pass\n```\n\n## Why this exists\n\nYou're building a production LLM app and your token bill is a line item:\n\n- RAG apps with large retrieved documents\n- Long-running conversational agents (LangChain / LangGraph / CrewAI)\n- Document-processing pipelines\n- Coding agents — Cursor-like / Claude-Code-like, with growing tool-call histories\n\nExisting options have gaps:\n\n- **Provider prompt caching** (Anthropic / OpenAI / Gemini) wins on stable prefixes — system prompts, tool definitions, retrieved-document pools. **It doesn't help with dynamic per-query content** (chat history, freshly retrieved docs, tool outputs). Compose with leanctx, don't choose between them.\n- **Naive truncation** drops the middle of the document, exactly where many answers live. The LongBench v2 numbers above show this concretely.\n- **Hosted compression APIs** (Compresr, Token Company) require sending your context to their servers. Closed-source models. leanctx is MIT-licensed, runs the model locally, and never makes outbound calls except to your existing provider.\n\n## Real numbers\n\n### Full LongBench v2 sweep — N=503, layered on a production compressor\n\nThe headline result. leanctx runs as a semantic pass **on top of** [ClawRouter](https://github.com/BlockRunAI/ClawRouter)'s seven structural compression layers, so the measured delta is what leanctx adds to a system that is *already* compressing. All 503 LongBench v2 questions (Tsinghua KEG, 8K–2M words), Claude Haiku 4.5 eval, temperature 0.1.\n\n| | Avg tokens / request | vs raw | vs Leg A |\n|---|---:|---:|---:|\n| Raw (uncompressed) | 27,865 | — | — |\n","type":"text"}]}
Thinking…

Important considerations:

  • LeanCTX is a local open-source Python library (MIT), not a hosted product. The user said "use the hosted product via SDK or API." There is no hosted product/API — it runs locally, compression happens on-device. But "SDK" — this is a pip library. The user said "If this product doesn't have a developer workflow you can work with, state that clearly and stop." A developer workflow does exist: pip install, use the SDK. But it's not "hosted." The instructions are "keep it light: use the hosted product via SDK or API. Don't start local service stacks." A pip library isn't a local service stack, so it's fine.

  • Pricing: free, MIT-licensed open source. There is no pricing itself. You continue to pay your own LLM provider fees (reduced by compression). This is the answer to "how does pricing work."

  • Building a simple example: a sample script. Real API calls require an API key. Without a key, can I demo compression with leanctx bench or Compressor directly? Lingua requires downloading about 1.2GB of model weights — that's a long download, which could be a blocker. Let me check whether the bench command works without weights. Actually, without lingua installed, it falls through to passthrough. Let me check if leanctx[lingua] is installable and what bench does. The user said not to wait for long-running commands, and a 1.2GB download is long. I can write an example script and run the offline bench to verify installation, without the lingua weights.

Let me check whether API keys (OPENAI_API_KEY / ANTHROPIC_API_KEY) are available in the environment.

$ env | grep -iE "openai|anthropic|gemini|google" | sed 's/=.*/=<set>/'; python -m leanctx.bench 2>/dev/null || leanctx bench list 2>&1 | head -20{"content":[{"text":"/bin/bash: line 1: leanctx: command not found\n","type":"text"}]}
$ curl -s https://raw.githubusercontent.com/jia-gao/leanctx/main/README.md | grep -iA6 -E "pricing|cost|free|licens" | head -40{"content":[{"text":"[![License](https://img.shields.io/pypi/l/leanctx)](LICENSE)\n\n**Drop-in prompt compression for production LLM applications.**\nCut your i…ctrl+o
{"content":[{"text":"[![License](https://img.shields.io/pypi/l/leanctx)](LICENSE)\n\n**Drop-in prompt compression for production LLM applications.**\nCut your input-token bill by 10–40%,without sacrificing accuracy.\n\n```python\n# before\n--\nOn the full **[LongBench v2](https://longbench2.github.io/)** set (N=503), layered on top of a compressor that is already running, leanctx removes **an extra 18.7 % of tokens** — rising to **36.7 % on prose-heavy traffic** — at a cost of 1.8 pp of accuracy. Every figure regenerates from [per-item records committed to this repo](#every-figure-here-regenerates-from-committed-data). Open-source models, runs locally, MIT-licensed. Your prompts and user data never leave your infrastructure by default.\n\n[Quickstart](#quickstart-60-seconds) · [What makes it different](#what-makes-it-different) · [Benchmarks](#real-numbers) · [Integrations](#integrations) · [How it works](#how-it-works)\n\n---\n\n## What makes it different\n--\n- **Hosted compression APIs** (Compresr, Token Company) require sending your context to their servers. Closed-source models. leanctx is MIT-licensed, runs the model locally, and never makes outbound calls except to your existing provider.\n\n## Real numbers\n\n### Full LongBench v2 sweep — N=503, layered on a production compressor\n\nThe headline result. leanctx runs as a semantic pass **on top of** [ClawRouter](https://github.com/BlockRunAI/ClawRouter)'s seven structural compression layers, so the measured delta is what leanctx adds to a system that is *already* compressing. All 503 LongBench v2 questions (Tsinghua KEG, 8K–2M words), Claude Haiku 4.5 eval, temperature 0.1.\n--\n**Compression costs accuracy, it does not add it.** The claim is that the cost is small and bounded — 1.8 pp overall, against the −2 pp go/no-go gate set for the integration. Because 54.2 % of Layer-8 input tokens route to verbatim, the compression actually applied to eligible content is **40.8 %**; the verbatim half contributes exactly 0 to both the savings and the accuracy delta by construction.\n\nSub-buckets are not uniform, and the report gives all of them: `short`/`lingua` is −17.6 pp (N=68) and Single-Document QA −11.5 pp (N=78), while Long Structured Data Understanding is +9.4 pp (N=32). Sidecar latency 47 ms p50 on GPU. At Sonnet input pricing the savings are ~$78 per 1,000 requests.\n\nFull report — per-bucket breakdowns by route × difficulty × length × domain, layer-by-layer contributions, cost model: [`benchmarks/clawrouter/full_long_bench_evaluation_result.md`](benchmarks/clawrouter/full_long_bench_evaluation_result.md).\n\n#### Independent execution and audit\n\nThis result was not produced by the maintainer. The benchmark was executed and audited by outside contributors ([@YingjingLu](https://github.com/YingjingLu), [@QianXiaoMoRan9](https://github.com/QianXiaoMoRan9)). An audit of an earlier draft showed that a headline \"+7.4 % on long context\" was **eval noise rather than a compression effect** — 7 of the 8 net improved items came from the verbatim subset, where the input was byte-identical (McNemar p = 0.143). Two methodology fixes followed ([#7](https://github.com/jia-gao/leanctx/pull/7)): Leg B now reuses Leg A's answer whenever the compressed context is byte-identical, so verbatim items contribute Δ = 0 by construction instead of decoder noise, and eval temperature dropped to 0.1. A separate correction forced a letter choice in the closed-book control, which had been scoring 6 % — below the 25 % random floor for 4-way multiple choice — and inflating apparent context lift.\n\nThe numbers above are post-correction. Discussion: [issue #3](https://github.com/jia-gao/leanctx/issues/3).\n--\n| Provider  | Model              | Compression | Latency   | Cost per call |\n|-----------|--------------------|:-----------:|:---------:|:-------------:|\n| Anthropic | `claude-haiku-4-5` | **41.6 %**  | 3.05 s    | ~$0.0016      |\n| OpenAI    | `gpt-4o-mini`      | **49.1 %**  | 6.42 s    | ~$0.0003      |\n","type":"text"}]}
Thinking…

Investigation findings:

  • LeanCTX is not a hosted product. It's an MIT-licensed local open-source Python library (a pip package). There's no hosted API, no pricing, and no external communication other than to your own LLM provider. It's explicitly positioned against hosted compression APIs.
  • The user's instructions: "Keep it light: use hosted products via SDK or API." This product has no hosted offering. However, a developer workflow usable via SDK does exist (pip install, wrap the client). Does this fall under "there is no developer workflow you can work with"? A workflow does exist — a pip SDK. The constraint is about not spinning up a local service stack — pip libraries are fine. But the quickstart with actual compression requires downloading 1.2GB of LLMLingua-2 weights (a time-consuming download), and API calls require a provider API key (I don't see one in the environment variables).

Can I build and run a simple example? Without an API key, the wrapper client can't make calls. Without the lingua weights, compression will pass through. That said, I can write an example script and demo the parts that work offline: classify/Compressor passthrough, tokens, stats. Let me check what works offline without weights — Verbatim compressor, classify, tokens count. Let me quickly test it.

$ cd /tmp && python - <<'EOF'{"content":[{"text":" File \"<stdin>\", line 2\n from leanctx.tokens import count_tokens if False else None\n ^^\…ctrl+o
cd /tmp && python - <<'EOF'
from leanctx import classify, ContentType, CompressionStats
from leanctx.tokens import count_tokens if False else None
import leanctx.tokens as t
print([x for x in dir(t) if not x.startswith('_')])
text = "The quick brown fox jumps over the lazy dog. " * 10
print(classify(text))
EOF

{"content":[{"text":"  File \"<stdin>\", line 2\n    from leanctx.tokens import count_tokens if False else None\n                                            ^^\nSyntaxError: invalid syntax\n\n\nCommand exited with code 1","type":"text"}],"details":{}}
Thinking…

Let's investigate the offline features (classifier, token counter, compressor), and then build a sample.

$ cd /tmp && python - <<'EOF'{"content":[{"text":"tokens: ['Any', 'annotations', 'count_message_tokens', 'count_tokens', 'get_text_content', 'lru_cache']\nclassifier: ['Any', 'ContentType',…ctrl+o
cd /tmp && python - <<'EOF'
import leanctx.tokens as t, leanctx.classifier as c
from leanctx.compressors import Verbatim, Lingua, ContentType
print("tokens:", [x for x in dir(t) if not x.startswith('_')])
print("classifier:", [x for x in dir(c) if not x.startswith('_')])
print("ContentType:", list(ContentType))
EOF

{"content":[{"text":"tokens: ['Any', 'annotations', 'count_message_tokens', 'count_tokens', 'get_text_content', 'lru_cache']\nclassifier: ['Any', 'ContentType', 'RepeatTracker', 'annotations', 'classify', 'get_text_content', 'hashlib']\nContentType: [<ContentType.UNKNOWN: 'unknown'>, <ContentType.PROSE: 'prose'>, <ContentType.CODE: 'code'>, <ContentType.ERROR: 'error'>, <ContentType.REPEAT: 'repeat'>, <ContentType.LONG_IMPORTANT: 'long_important'>]\n","type":"text"}]}
Thinking…

Let's test the compressor offline (without weights, lingua falls back to passthrough), write a simple sample script, and verify that it can run without an API key. Since there's no API key in the environment (grep didn't find anything), the wrapper examples can't actually be run against a real provider. I can build a sample that demonstrates the offline pipeline (classify + compress + stats) and show a drop-in client snippet that would work if an API key were present.

$ cd /tmp && python - <<'EOF'{"content":[{"text":"(self, *args, **kwargs)\n(self, messages: 'list[dict[str, Any]]') -> 'tuple[list[dict[str, Any]], CompressionStats]'\n(self, messages: 'lis…ctrl+o
cd /tmp && python - <<'EOF'
from leanctx.compressors import Compressor, Lingua, Verbatim
import inspect
print(inspect.signature(Compressor.__init__))
print(inspect.signature(Compressor.compress) if hasattr(Compressor,'compress') else [m for m in dir(Compressor) if not m.startswith('_')])
print(inspect.signature(Lingua.compress) if hasattr(Lingua,'compress') else '')
EOF

{"content":[{"text":"(self, *args, **kwargs)\n(self, messages: 'list[dict[str, Any]]') -> 'tuple[list[dict[str, Any]], CompressionStats]'\n(self, messages: 'list[dict[str, Any]]') -> 'tuple[list[dict[str, Any]], CompressionStats]'\n","type":"text"}]}
Thinking…

Let's build a simple sample. Since no API keys are set (environment variables are empty), calls to the live provider won't work. However, the offline pipeline (classify + compress via Compressor using passthrough/verbatim, or lingua if weights are available) can be executed. Lingua requires a 1.2GB weight download, which falls under a "long-running" process, so skip it. Without the lingua extra, fall back to passthrough. Check whether lingua's dependencies are installed.

$ pip list 2>/dev/null | grep -iE "llmlingua|torch|transformers|anthropic|openai"{"content":[{"text":"anthropic 0.122.0\nllama-index-embeddings-openai 0.6.0\nllama-index-llms-openai …ctrl+o
{"content":[{"text":"anthropic                                    0.122.0\nllama-index-embeddings-openai                0.6.0\nllama-index-llms-openai                      0.7.10\nopenai                                       2.54.0\nopenai-agents                                0.20.0\n","type":"text"}]}
Thinking…

No API key is set, and lingua/torch aren't installed either (the Lingua compressor needs roughly 1.2GB of model weights — that's a long download; the user said to keep it light and not to wait for long-running commands). The offline parts can be demoed: classify + Verbatim/pass-through compression, and the bench module. Let me try writing a small example in /sandbox/repo that works fully offline: classification, verbatim compression, token counting, and (optionally) a drop-in OpenAI client snippet that would work if an API key were available.

Let's test Compressor with the verbatim route.

There's no API key in the environment variables, and the Lingua compressor requires a 1.2 GB model weight download — so, I'll create a sample that runs fully offline (classification + compression pipeline + statistics), and additionally attach a snippet of a drop-in client for when a key does become available. First, let's confirm that the offline pipeline works.

$ cd /tmp && python - <<'EOF'{"content":[{"text":"(self, *args, **kwargs)\n['apply']\ntokens before: 2531\nTraceback (most recent call last):\n File \"<stdin>\", line 11, in <module>\n Fi…ctrl+o
cd /tmp && python - <<'EOF'
from leanctx import Compressor, Strategy, classify, ContentType
from leanctx.tokens import count_message_tokens
import inspect
print(inspect.signature(Strategy.__init__) if hasattr(Strategy,'__init__') else '')
print([m for m in dir(Strategy) if not m.startswith('_')])
msgs = [
    {"role": "system", "content": "You are a helpful assistant. Always answer in JSON."},
    {"role": "user", "content": "Here is a document. " * 500 + "\n\n```python\ndef f(x):\n    return x*2\n```\nSummarize."},
]
print("tokens before:", count_message_tokens(msgs))
c = Compressor()
out, stats = c.compress(msgs)
print("tokens after:", count_message_tokens(out))
print("stats:", stats)
for m in msgs[1:]:
    print(classify(m["content"]))
EOF

{"content":[{"text":"(self, *args, **kwargs)\n['apply']\ntokens before: 2531\nTraceback (most recent call last):\n  File \"<stdin>\", line 11, in <module>\n  File \"/usr/lib/python3.12/typing.py\", line 1742, in _no_init_or_replace_init\n    raise TypeError('Protocols cannot be instantiated')\nTypeError: Protocols cannot be instantiated\n\n\nCommand exited with code 1","type":"text"}],"details":{}}
Thinking…

Compressor is a Protocol. Use Lingua or Verbatim instead. Lingua requires the llmlingua package (not installed)—it probably falls back to passthrough. Let me try Verbatim and Lingua.

$ cd /tmp && python - <<'EOF'{"content":[{"text":"tokens before: 2531\nVerbatim after: 2531 stats: CompressionStats(input_tokens=2531, output_tokens=2531, ratio=1.0, method='verbatim', cost…ctrl+o
cd /tmp && python - <<'EOF'
from leanctx import Lingua, Verbatim, classify
from leanctx.tokens import count_message_tokens
msgs = [
    {"role": "system", "content": "You are a helpful assistant. Always answer in JSON."},
    {"role": "user", "content": "Here is a document. " * 500 + "\n\n```python\ndef f(x):\n    return x*2\n```\nSummarize."},
]
print("tokens before:", count_message_tokens(msgs))
for C in (Verbatim, Lingua):
    c = C()
    out, stats = c.compress(msgs)
    print(C.__name__, "after:", count_message_tokens(out), "stats:", stats)
print("classify prose:", classify("Here is a document. " * 500))
print("classify code:", classify("```python\ndef f(x):\n    return x*2\n```"))
EOF

{"content":[{"text":"tokens before: 2531\nVerbatim after: 2531 stats: CompressionStats(input_tokens=2531, output_tokens=2531, ratio=1.0, method='verbatim', cost_usd=0.0)\nTraceback (most recent call last):\n  File \"/opt/freestyle/python/lib/python3.12/site-packages/leanctx/compressors/lingua.py\", line 268, in _load\n    import llmlingua\nModuleNotFoundError: No module named 'llmlingua'\n\nThe above exception was the direct cause of the following exception:\n\nTraceback (most recent call last):\n  File \"<stdin>\", line 10, in <module>\n  File \"/opt/freestyle/python/lib/python3.12/site-packages/leanctx/compressors/lingua.py\", line 117, in compress\n    stats = self._compress_inner(messages)\n            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"/opt/freestyle/python/lib/python3.12/site-packages/leanctx/compressors/lingua.py\", line 136, in _compress_inner\n    new_msg, in_tok, out_tok = self._compress_message(msg)\n                               ^^^^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"/opt/freestyle/python/lib/python3.12/site-packages/leanctx/compressors/lingua.py\", line 175, in _compress_message\n    compressed, in_tok, out_tok = self._compress_text(content)\n                                  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"/opt/freestyle/python/lib/python3.12/site-packages/leanctx/compressors/lingua.py\", line 253, in _compress_text\n    compressor = self._load()\n                 ^^^^^^^^^^^^\n  File \"/opt/freestyle/python/lib/python3.12/site-packages/leanctx/compressors/lingua.py\", line 270, in _load\n    raise ImportError(_LINGUA_INSTALL_HINT) from e\nImportError: The 'llmlingua' package is required for leanctx.Lingua. Install with: pip install 'leanctx[lingua]'\n\n\nCommand exited with code 1","type":"text"}],"details":{}}
Thinking…

Lingua requires installing llmlingua and downloading approximately 1.2GB of weights, which is a time-consuming download. The user instructed to keep it lightweight and avoid long-running commands. Installing llmlingua itself can be done quickly with pip, but the model weights download (1.2GB) occurs on the first invocation. This may be too heavy. Options: either build an example without actual compression (in pass-through mode), or try installing the lingua extra. First, let's quickly check the size of pip install llmlingua with a dry-run. In practice, llmlingua pulls in torch/transformers, which becomes a several-GB-scale download. This falls under the "long-running" territory. It would be better to provide examples that work offline: classification, verbatim/pass-through, stats, and a drop-in client example (which requires an API key). Also, bench can run offline too, so maybe it can be demonstrated with that. Let's check the bench module.

Lingua mode requires llmlingua plus about 1.2GB of model weights (torch/transformers), and since that counts as a long download, we'll keep it lightweight and skip it. Let's check out the offline bench that's included in the package:

$ cd /tmp && python -m leanctx.bench list 2>&1 | head -20; ls /opt/freestyle/python/lib/python3.12/site-packages/leanctx/bench{"content":[{"text":"/opt/freestyle/python/bin/python3: No module named leanctx.bench.__main__; 'leanctx.bench' is a package and cannot be directly executed\n__…ctrl+o
{"content":[{"text":"/opt/freestyle/python/bin/python3: No module named leanctx.bench.__main__; 'leanctx.bench' is a package and cannot be directly executed\n__init__.py\n__pycache__\ncli.py\nrunners\nscenarios.py\nschema.py\nworkloads.py\n","type":"text"}]}
Thinking…

Lingua requires installing llmlingua (which entails a large torch download, roughly 1.2GB of weights). This falls under what the user calls a "long-running" operation. The user said, "Keep it lightweight... don't wait for long-running commands." So, I'll build an example that works offline: pass-through/verbatim mode using the drop-in client, plus a bench CLI. Check whether the leanctx CLI exists as an entry point? "leanctx: command not found" — the script might not be installed. Let me try python -m leanctx.bench.cli.

$ cd /tmp && python -c "from leanctx.bench.cli import main; main()" list 2>&1 | head; python -m leanctx.bench.cli list 2>&1 | head -15{"content":[{"text":" agent-structural Structural-integrity invariants on the agent workload. [missing extras: lingua]\n anthropic-e2e F…ctrl+o
{"content":[{"text":"  agent-structural          Structural-integrity invariants on the agent workload.  [missing extras: lingua]\n  anthropic-e2e             Full leanctx.Anthropic stack with respx-mocked Anthropic API.  [missing extras: bench]\n  lingua-local              Offline LLMLingua-2 compression on a single workload.  [missing extras: lingua]\n  selfllm-anthropic         SelfLLM compression via Anthropic (Haiku) live API.  [missing env: ANTHROPIC_API_KEY]\n  selfllm-openai            SelfLLM compression via OpenAI (gpt-4o-mini) live API.  [missing env: OPENAI_API_KEY]\n  selfllm-gemini            SelfLLM compression via Gemini (2.5-flash) live API.  [missing env: GEMINI_API_KEY|GOOGLE_API_KEY]\n  agent-structural          Structural-integrity invariants on the agent workload.  [missing extras: lingua]\n  anthropic-e2e             Full leanctx.Anthropic stack with respx-mocked Anthropic API.  [missing extras: bench]\n  lingua-local              Offline LLMLingua-2 compression on a single workload.  [missing extras: lingua]\n  selfllm-anthropic         SelfLLM compression via Anthropic (Haiku) live API.  [missing env: ANTHROPIC_API_KEY]\n  selfllm-openai            SelfLLM compression via OpenAI (gpt-4o-mini) live API.  [missing env: OPENAI_API_KEY]\n  selfllm-gemini            SelfLLM compression via Gemini (2.5-flash) live API.  [missing env: GEMINI_API_KEY|GOOGLE_API_KEY]\n","type":"text"}]}
Thinking…

I've got a clear grasp of the situation. LeanCTX is not a hosted product — it's an MIT-licensed local Python library (an SDK that wraps the anthropic/openai/gemini clients). Regarding pricing: it's free/open source, and there's no hosted API to pay for. The "pricing" angle is that it reduces your provider's token costs.

The user's question is "how does pricing work" — answer: the software itself has no pricing, it's MIT open source, you only pay your LLM provider, and leanctx reduces input tokens.

"Is this product easy for you to manage" — yes: pure pip SDK, no services, no accounts, no keys needed for local compression. However, full compression requires about 1.2GB of LLMLingua weights (long download — avoid). The quickstart doesn't require a local service stack.

Building a simple example: a runnable offline sample that demonstrates classification + verbatim passthrough + stats, and a drop-in OpenAI/Anthropic client sample (keys required — not available). I'll write a sample file at /sandbox/repo and run the offline portion.

I've explored LeanCTX enough to build the example. Let me write it and verify the offline parts run.

write /sandbox/repo/leanctx_example.py{"content":[{"text":"Successfully wrote 2986 bytes to /sandbox/repo/leanctx_example.py","type":"text"}]}
Thinking…

Let me test this to make sure it works as expected.

$ cd /sandbox/repo && python leanctx_example.py{"content":[{"text":"== classify segments (loss-tolerance routing) ==\nTraceback (most recent call last):\n File \"/sandbox/repo/leanctx_example.py\", line 36,…ctrl+o
{"content":[{"text":"== classify segments (loss-tolerance routing) ==\nTraceback (most recent call last):\n  File \"/sandbox/repo/leanctx_example.py\", line 36, in <module>\n    print(f\"  prose doc  -> {classify(LONG_DOC).value}   (high tolerance: compressible)\")\n                             ^^^^^^^^^^^^^^^^^^\n  File \"/opt/freestyle/python/lib/python3.12/site-packages/leanctx/classifier.py\", line 70, in classify\n    text = get_text_content(message)\n           ^^^^^^^^^^^^^^^^^^^^^^^^^\n  File \"/opt/freestyle/python/lib/python3.12/site-packages/leanctx/_content.py\", line 42, in get_text_content\n    return _extract_text(message.get(\"content\"))\n                         ^^^^^^^^^^^\nAttributeError: 'str' object has no attribute 'get'\n\n\nCommand exited with code 1","type":"text"}],"details":{}}
Thinking…

classify expects a message dict. Please fix it.

edit /sandbox/repo/leanctx_example.py{"content":[{"text":"Successfully replaced 1 block(s) in /sandbox/repo/leanctx_example.py.","type":"text"}],"details":{"diff":" ...\n 32 {\"role\": \"use…ctrl+o
{"content":[{"text":"Successfully replaced 1 block(s) in /sandbox/repo/leanctx_example.py.","type":"text"}],"details":{"diff":"    ...\n 32     {\"role\": \"user\", \"content\": f\"Context:\\n{LONG_DOC}\\n\\n{CODE_BLOCK}\\n\\nSummarize and compute discount(100, 20).\"},\n 33 ]\n 34 \n 35 print(\"== classify segments (loss-tolerance routing) ==\")\n-36 print(f\"  prose doc  -> {classify(LONG_DOC).value}   (high tolerance: compressible)\")\n-37 print(f\"  code block -> {classify(CODE_BLOCK).value}   (zero tolerance: verbatim)\")\n+36 print(f\"  prose doc  -> {classify({'role': 'user', 'content': LONG_DOC}).value}   (high tolerance: compressible)\")\n+37 print(f\"  code block -> {classify({'role': 'user', 'content': CODE_BLOCK}).value}   (zero tolerance: verbatim)\")\n 38 \n 39 before = count_message_tokens(messages)\n 40 \n 41 # Verbatim compressor: byte-for-byte passthrough (baseline; the Lingua\n    ...","patch":"--- /sandbox/repo/leanctx_example.py\n+++ /sandbox/repo/leanctx_example.py\n@@ -32,10 +32,10 @@\n     {\"role\": \"user\", \"content\": f\"Context:\\n{LONG_DOC}\\n\\n{CODE_BLOCK}\\n\\nSummarize and compute discount(100, 20).\"},\n ]\n \n print(\"== classify segments (loss-tolerance routing) ==\")\n-print(f\"  prose doc  -> {classify(LONG_DOC).value}   (high tolerance: compressible)\")\n-print(f\"  code block -> {classify(CODE_BLOCK).value}   (zero tolerance: verbatim)\")\n+print(f\"  prose doc  -> {classify({'role': 'user', 'content': LONG_DOC}).value}   (high tolerance: compressible)\")\n+print(f\"  code block -> {classify({'role': 'user', 'content': CODE_BLOCK}).value}   (zero tolerance: verbatim)\")\n \n before = count_message_tokens(messages)\n \n # Verbatim compressor: byte-for-byte passthrough (baseline; the Lingua\n","firstChangedLine":36}}
Thinking…

Let me test this to make sure it works as expected.

$ cd /sandbox/repo && python leanctx_example.py{"content":[{"text":"== classify segments (loss-tolerance routing) ==\n prose doc -> prose (high tolerance: compressible)\n code block -> code (zero tole…ctrl+o
{"content":[{"text":"== classify segments (loss-tolerance routing) ==\n  prose doc  -> prose   (high tolerance: compressible)\n  code block -> code   (zero tolerance: verbatim)\n\n== compression stats ==\n  tokens before : 875\n  tokens after  : 875\n  stats         : CompressionStats(input_tokens=875, output_tokens=875, ratio=1.0, method='verbatim', cost_usd=0.0)\n","type":"text"}]}
Thinking…

The example runs. Now let's write the answer: how the pricing mechanism works, whether management is easy, and points to watch out for.

Pricing: LeanCTX itself is free — MIT-licensed open source, runs locally, no hosted service or API to pay for. You only pay your LLM provider (OpenAI/Anthropic/Gemini), and leanctx reduces input tokens (claimed to be 10–40%). The optional "self-LLM" mode incurs a small cost for cheap model calls (e.g., about $0.0003 per call with gpt-4o-mini).

Important note: the user said "use the hosted product via its SDK or API." LeanCTX is not a hosted product — it's a local library, and no hosted service exists. I should state this clearly. However, since a developer workflow (pip SDK) exists, it doesn't fall under "no developer workflow." That said, since it's local-only, I should point out that the premise of "hosted product" doesn't apply. The SDK runs entirely locally, and no service is required.

Ease of management: Yes — a single pip package, drop-in wrapper, offline pipeline is verifiable, no keys required for the basics. Caveats: full compression requires leanctx[lingua], which downloads about 1.2GB of model weights (I skipped this to keep things lightweight), and live calls require provider API keys (not configured in this environment).

Done. Here's the summary:

What LeanCTX actually is (one plain caveat first)

LeanCTX is not a hosted product — it's an MIT-licensed Python SDK (pip install leanctx) that runs compression locally and wraps the official OpenAI/Anthropic/Gemini clients. There's no hosted API or service to call. It does have a real developer workflow I could act on, so I proceeded — everything stays in-process, no local service stacks, nothing long-running.

The example

Created /sandbox/repo/leanctx_example.py — runs offline with no API keys:

  • Part 1 (verified running): the core loss-tolerance pipeline — classify() routes prose as compressible and code as zero-tolerance/verbatim, then a compressor pass with CompressionStats telemetry. Output confirmed: prose → prose, code → code, stats reported correctly.
  • Part 2 (commented): the drop-in client pattern — from leanctx import OpenAI with a leanctx_config (mode, token threshold, prose→Lingua routing), plus usage.leanctx_tokens_saved / leanctx_ratio telemetry on responses. Needs a provider API key, which isn't set in this environment.

How pricing works

  • The SDK itself is free — MIT-licensed, runs on-device, no subscriptions, no metered hosted API, no outbound calls except to your existing LLM provider.
  • You pay only your provider's normal token bills — leanctx's value is reducing them (claims ~10–40% input-token savings; ~$78 per 1,000 requests at Sonnet input pricing per their benchmark).
  • One indirect cost: the opt-in "self-LLM" mode uses a cheap model (e.g., gpt-4o-mini, ~$0.0003/call) to do compression, so it adds a small per-request charge.

Will it be easy for me to manage?

Yes, mostly. Single pip package, clean typed SDK, works offline for verification, drop-in client wrappers, and a built-in bench CLI for invariant checks. Two friction points:

  1. Real compression needs pip install 'leanctx[lingua]', and first use downloads ~1.2 GB of model weights — I deliberately skipped that per your "stay light" constraint, so the demo uses the verbatim/passthrough compressor.
  2. Live end-to-end runs need provider API keys (OPENAI_API_KEY etc.), which aren't present here — that's the only thing blocking a full compressed request demo.

AX Fundamentals

23 checks. Open an area for details.

This check is a quick preview. Test more tasks and models with Gauge Agents.

Test your product with Gauge Agents

AX Check by Gauge