by NiceEval
build eval for your agent in 10 mins
# Add to your Claude Code skills
git clone https://github.com/NiceEval/NiceEvalLast scanned: 8/23/2026
{
"issues": [
{
"type": "npm-audit",
"message": "brace-expansion: brace-expansion: DoS via unbounded intermediate arrays, bypassing the CVE-2026-14257 mitigation",
"severity": "high"
},
{
"type": "npm-audit",
"message": "nx: Vulnerability found",
"severity": "high"
}
],
"status": "WARNING",
"scannedAt": "2026-08-23T04:35:53.512Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}See how NiceEval compares with popular alternatives.
NiceEval is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by NiceEval. build eval for your agent in 10 mins. It has 156 GitHub stars.
NiceEval returned warnings in SkillsLLM's automated security scan. It has no critical vulnerabilities, but review the flagged issues in the Security Report section before adding it to your workflow.
Clone the repository with "git clone https://github.com/NiceEval/NiceEval" and add it to your Claude Code skills directory (see the Installation section above).
NiceEval is primarily written in TypeScript. It is open-source under NiceEval on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh NiceEval against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
Write evals like unit tests, compare agents like experiments
中文 | Deutsch | Español | français | 日本語 | 한국어 | Português | Русский
You revised a prompt, switched models, wrote a new Skill for Claude Code, or tuned the prompts for NPCs in a game. Did it actually get better?
Most of the time, the answer comes from trying it a few times and going by feel. But LLM outputs are unpredictable: the same question may get a correct answer now and a wrong one next time. Three successful attempts don't prove that a change works, or that it hasn't broken something else.
NiceEval replaces guesswork with evidence. Its main focus is agents, but it can evaluate any LLM-powered application. You describe “what counts as correct” in TypeScript: which tool to call, what a reply should contain, whether tests pass after a code change, or whether a game world remains coherent after an interaction. NiceEval connects to the system under test, runs it repeatedly, scores it, and keeps each run's conversations, tool calls, file changes, timing, and cost for comparison and investigation.
Writing evals feels like writing unit tests; using them feels more like running experiments:
Everything runs on your own machine and in your CI. No account required.
Check a weather assistant: when asked about the weather, it should actually call get_weather instead of making up an answer.
// evals/weather-tool.eval.ts
import { defineEval, defineJudge } from "niceeval";
import { includes, jsonMatch, pattern, toolMatch } from "niceeval/expect";
const groundedAnswer = defineJudge({
name: "grounded-weather-answer",
rubric: "Does the assistant answer using the tool's weather data rather than refusing or being vague?",
});
export default defineEval({
description: "Call the weather tool and answer using its result",
judge: groundedAnswer,
async test(t) {
const turn = await t.send("What's the weather in Beijing today?");
turn.succeeded();
// Check deterministic facts with deterministic rules
turn.calledTool(toolMatch("get_weather", { input: jsonMatch({ city: "Beijing" }) }));
t.check(turn.message, pattern(/°C|temperature|sunny|cloudy|rain/));
// The second turn should retain the context
const second = await t.send("What about Shanghai tomorrow?");
t.check(second.message, includes("Shanghai"));
// Let a judge model assess open-ended quality
t.judge({ question: turn.input, answer: turn.message }, groundedAnswer).gate(0.7);
},
});
The Experiment specifies “which agent and which model to run,” separately from the eval cases. You can therefore use the same cases to compare two models or two prompt versions directly:
// experiments/local.ts
import { defineExperiment } from "niceeval";
import { webAgent } from "../agents/web-agent"; // Your own Adapter, a few dozen lines
export default defineExperiment({
agent: webAgent({ baseUrl: "http://127.0.0.1:5188" }),
model: "gpt-5.5",
});
pnpm exec niceeval exp local weather-tool # Run only weather-tool with the local experiment
pnpm exec niceeval show # See results in the terminal, failures first
pnpm exec niceeval view # Browse conversations and tool calls in the browser
A complete runnable project is available in examples/zh/ai-sdk/.
Your own AI applications. Whether built with AI SDK, LangGraph, Pi, or your own agent loop, and regardless of implementation language, they just need an interface you can call: HTTP, WebSocket, or an SDK. Write an Adapter that sends requests and translates replies into events NiceEval can read, then assert on reply content, tool calls, structured output, and usage.
Coding agents and their extensions. NiceEval runs agents such as Claude Code, Codex, and OpenCode in Docker or a cloud Sandbox, gives them a real repository and a task, and scores the result using the project's own tests and file changes. Use it to answer: “Does this Skill, Plugin, or memory actually help the agent do better work?”
Any AI application, including LLM games. The system under test doesn't have to be a conversational agent. LLM-powered games, AI social apps, and generative workflows often expose operations such as posting, replying, or NPC actions instead of back-and-forth messages. Use defineAdapter to expose those operations directly to eval cases. Call typed methods such as t.post(...) and t.reply(...), check structured results and world state, and leave open-ended quality to a Judge.
// evals/social-journey.eval.ts — the system under test is an AI-powered social app
export default x.defineEval({
description: "Post and reply, get responses from AI characters, and keep the social world coherent",
async test(t) {
const initial = await t.visitDiscoveryPage();
const post = await t.post({ intent: "Invite everyone to photograph the city at night", withImage: true });
t.check(post, authoredPost(initial.viewerId)).label("The post belongs to the current player");
const reply = await t.reply({ postId: post.id, intent: "Add the meeting point at the riverside path entrance" });
const responses = await t.waitForReplies(reply.id);
t.check(responses.length, greaterThan(0)).label("AI characters respond");
t.check(worldMaterial(await t.refreshFeed()), coherentSocialWorld()).label("Social relationships remain intact after refresh");
},
});
See examples/zh/llm-x/ for a complete AI-powered social app project. Its evals cover posting, replying, AI character responses, refreshing the timeline, and generating accompanying images.
DeepEval is a mature Python evaluation framework with a rich metric library. NiceEval makes different tradeoffs:
Observability platforms such as LangFuse and Braintrust answer “what happened in production”; evals answer “is this behavior good enough?” NiceEval focuses on the latter and the local development loop of writing evals, running them, inspecting results, and improving the agent. They can coexist: keep using those platforms for production traces and send NiceEval results to Braintrust.
The fastest way to get started is to ask the coding agent you already use to integrate NiceEval. Send it this:
Read https://niceeval.com/INIT.md, install and integrate niceeval into the current repository, and run the first eval case end to end.
To do it yourself, follow the quickstart and write three files. You can see your first result in about ten minutes.
This project was inspired by the projects below, and some code was written by AI learning from them:
Thanks to Linux.do for their support and feedback during the project's early development.