by Sev7eEn7
dsh-sieve: context engineering & token optimization plugin for DeepSeek Harness (DSH) — tool output filtering, context pruning, progressive skill disclosure. 36% smaller payload in offline replay. DSH 上下文管理与 token 优化插件。
# Add to your Claude Code skills
git clone https://github.com/Sev7eEn7/dsh-sieveSee how dsh-sieve compares with popular alternatives.
dsh-sieve is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by Sev7eEn7. dsh-sieve: context engineering & token optimization plugin for DeepSeek Harness (DSH) — tool output filtering, context pruning, progressive skill disclosure. 36% smaller payload in offline replay. DSH 上下文管理与 token 优化插件。. It has 51 GitHub stars.
dsh-sieve's catalog security scan is still queued. You can run an instant dependency and prompt-injection check now with the "Scan for vulnerabilities" button above.
Clone the repository with "git clone https://github.com/Sev7eEn7/dsh-sieve" and add it to your Claude Code skills directory (see the Installation section above).
dsh-sieve is primarily written in TypeScript. It is open-source under Sev7eEn7 on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh dsh-sieve against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
The deep catalog scan for this skill is still queued. Run an instant dependency check now instead.
English | 简体中文
LLM agent context engineering and token efficiency plugins for DeepSeek Harness (DSH)
sieve filters tool output, prunes older context and progressively discloses skills to reduce redundant input tokens in AI coding agents, improve context window utilization and provide observable usage data for LLM inference cost optimization. Originals are archived before rewriting and can be retrieved on demand.
Over 30% less request payload in offline replay, measured in characters: a 36.28% aggregate reduction across 100 public agent trajectories. See offline replay results.
Historical live model results: 21.2% lower main model input token usage across 80 paired runs on deepseek-flash, with cache hit rate moving from 93.4% to 92.9%. This experiment used the now-removed session-model judge path; the current Jev / Laya configuration has not been remeasured. See live model results and the cache FAQ.
Built around context engineering and token efficiency, sieve manages the context delivered to the model within the agent loop, reducing repeated input across multi-step tool use.
[!IMPORTANT] sieve works only with a judge model configured: Jev (hosted, needs an API key) or Laya (local, macOS on Apple Silicon). sieve never uses the session's model in its place. With neither configured, the plugin loads but changes no context. See judge models.
In an agent loop the main model rereads its whole context at every step. What makes requests large is rarely the user's input. It is:
All of this is billed as input tokens, fills the context window and dilutes the model's attention. sieve adds a judgment kernel on DSH's public extension points that admits, forgets and progressively discloses these three kinds of content, so the main model sees a shorter, more focused context at every step.
The hard part is deciding what can go. Fixed rules have to keep any output format they have not seen; asking the main model to decide costs its tokens and a full round of reasoning. sieve hands this step to a dedicated judge model that answers structured yes/no and single-choice questions with probabilities, and the kernel rewrites or keeps the original by threshold. In offline replay, the median latency of one Jev judgment was about 0.3 seconds.
For AI coding agents running on DeepSeek Harness, especially tasks with verbose test logs, frequent tool calls, long sessions or large skill catalogs.
| Optimization area | How sieve handles it |
|---|---|
| Input token reduction | Filters redundant tool content before it enters the session, reducing input resent on later steps |
| Tool output filtering | Handles build logs, test reports and command output by category, protecting failure details and test summaries |
| Context compaction / pruning | Replaces irrelevant older tool results in batches, retaining archive pointers for on-demand retrieval |
| Progressive skill disclosure | Shows skills by task relevance, reducing context window usage from unrelated descriptions |
| Prompt cache awareness | Processes new results before they are recorded and batches historical pruning to control prefix rebuild frequency |
Install the plugin (the web profile also gets the panel):
dsh plugin --profile web add dsh-sieve@0.2.0 dsh-sieve-web@0.2.0
Configure a judge model, one of:
TYPESAFE_API_KEY environment variable;judge.type to laya in the profile, see Using Laya.Restart dsh web and reload the page, then enter /sieve status in a session and check that the "判断模型" (judge model) line does not say "未配置" (not configured).
For other ways to install (GitHub Release, building from source) see Installation.
flowchart LR
subgraph Loop["DSH Agent Loop"]
T[Tool execution] -->|post-execute| A
P[agent/pre-step] --> F
S[Skill catalog injection] --> D
end
subgraph Sieve["dsh-sieve"]
A["tool.admission<br/>tool.admission.test-log"]
F[context.forget]
D[skills.disclosure]
K(("Judgment kernel<br/>Judge Engine"))
A <--> K
F <--> K
D <--> K
end
K <-->|"System One / local HTTP"| J[("Judge model<br/>Jev · Laya")]
A & F -->|originals| SP[(spill archive)]
K -->|decision records| L[(ctx.storage ledger)]
L --> W["/sieve command · Web panel"]
A & F & D -->|reduced context| M[Main model request]
| Component | Role |
|---|---|
| Judgment kernel | One decision engine: builds the questions, calls the judge model, validates the structured answers and decides by confidence threshold whether to rewrite; a timeout or failure always keeps the original |
tool.admission |
Admission control for tool results before they enter the session; handles long output in tiers and rewrites only content, never the tool call |
tool.admission.test-log |
A dedicated admission path for test runs that protects failures, summaries and anything it does not recognize |
context.forget |
Forgets older tool results in batches through DSH's native compaction/prune and result replacement; a batch is committed only when it saves enough, see the cache FAQ |
skills.disclosure |
Filters the first skill catalog by task relevance and announces skills that become relevant in later turns; hidden skills can still be loaded directly |
| Ledger | One ctx.storage document per session with each judgment's result, usage and estimated saving |
dsh-sieve-web |
DSH Web sidebar panel: the session's estimated input token reduction and context reduction in characters, the judge model in use, and Jev key configuration |
What rules can handle, deterministic rules handle. What they cannot settle goes to the judge model. A decision point turns the content into structured questions (yes/no, single choice); the judge model answers each with a probability or an option, and the kernel rewrites or keeps the original by threshold.
flowchart LR
D[Decision point] --> R{Judge model configured?}
R -->|no| N[Change nothing<br/>rules do not run either]
R -->|yes| Q[Structured questions<br/>redacted]
Q --> JV["Jev<br/>System One protocol"]
Q --> LY["Laya<br/>local sidecar"]
JV & LY -->|probabilities / options| P{Threshold}
P -->|passes| AP[Rewrite context]
P -->|fails / timeout / error| KP[Keep original]
| Judge model | Description | Needs |
|---|---|---|
| Jev | TypeSafe's structured judgment model. It answers yes/no and single-choice questions natively over the System One protocol, with a probability for each answer. Reached through TypeSafe's own endpoint or OpenRouter, billed by TypeSafe (official pricing) | A TypeSafe or OpenRouter API key |
| Laya | An open-weight typed decision model (mmBERT-base, 322M) running on Core ML on this machine: requests never leave it and cost nothing. Its window is only 1024 tokens, so long states are cut, and it is clearly weaker than Jev on complex questions | macOS on Apple Silicon, with the local Laya server running |
No session model. Judgments go only to Jev or Laya. sieve never hands a judgment to the session's main model and has no setting for it; a profile from an older version with judge.type: llm, judge.type: off, judge.provider or routes fails to load.
No judge model, no changes. The default judge.type: auto looks for a Jev key before each decision. Without one, all four decisions let everything through: nothing is folded, forgotten or filtered, and nothing is recorded; /sieve status and the web panel say no judge model is configured. A saved key takes effect from the next decision, without a restart.
Judgments stay out of the conversation. A judge call is a side request: it is not written to the session log, does not enter the main model's context and does not change the main model's request prefix. What goes to the judge is redacted first: only content that is certainly a credential is replaced, namely credential-shaped tokens, assignments w