Dive Into LangGraph
luochang212/dive-into-langgraph
A Chinese-language guide and reference for building agents with LangGraph 1.0, from a first ReAct agent through middleware, memory, MCP, RAG and web search.
Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.
$ npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentailor/fullstack-langgraph-nextjs-agent agent-eval-cases --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentailor/fullstack-langgraph-nextjs-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/agent-eval-cases .claude/skills/agent-eval-cases && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-eval-cases" agent skill from https://github.com/agentailor/fullstack-langgraph-nextjs-agent/tree/main/.agents/skills/agent-eval-cases into .claude/skills/agent-eval-cases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval-cases", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentailor/fullstack-langgraph-nextjs-agent/tree/main/.agents/skills/agent-eval-casesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentailor/fullstack-langgraph-nextjs-agent agent-eval-cases --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentailor/fullstack-langgraph-nextjs-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/agent-eval-cases .agents/skills/agent-eval-cases && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-eval-cases" agent skill from https://github.com/agentailor/fullstack-langgraph-nextjs-agent/tree/main/.agents/skills/agent-eval-cases into .agents/skills/agent-eval-cases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval-cases", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentailor/fullstack-langgraph-nextjs-agent agent-eval-cases --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentailor/fullstack-langgraph-nextjs-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/agent-eval-cases .cursor/skills/agent-eval-cases && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-eval-cases" agent skill from https://github.com/agentailor/fullstack-langgraph-nextjs-agent/tree/main/.agents/skills/agent-eval-cases into .cursor/skills/agent-eval-cases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval-cases", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentailor/fullstack-langgraph-nextjs-agent.git --path .agents/skills/agent-eval-cases--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentailor/fullstack-langgraph-nextjs-agent agent-eval-cases --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentailor/fullstack-langgraph-nextjs-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/agent-eval-cases .gemini/skills/agent-eval-cases && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-eval-cases" agent skill from https://github.com/agentailor/fullstack-langgraph-nextjs-agent/tree/main/.agents/skills/agent-eval-cases into .gemini/skills/agent-eval-cases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval-cases", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentailor/fullstack-langgraph-nextjs-agent agent-eval-casesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentailor/fullstack-langgraph-nextjs-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/agent-eval-cases .github/skills/agent-eval-cases && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-eval-cases" agent skill from https://github.com/agentailor/fullstack-langgraph-nextjs-agent/tree/main/.agents/skills/agent-eval-cases into .github/skills/agent-eval-cases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval-cases", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentailor/fullstack-langgraph-nextjs-agent agent-eval-cases --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentailor/fullstack-langgraph-nextjs-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/agent-eval-cases .opencode/skills/agent-eval-cases && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-eval-cases" agent skill from https://github.com/agentailor/fullstack-langgraph-nextjs-agent/tree/main/.agents/skills/agent-eval-cases into .opencode/skills/agent-eval-cases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval-cases", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-eval-casesDecide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.
Agent Eval Cases is an agent skill from agentailor/fullstack-langgraph-nextjs-agent. Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. Use when writing a first eval suite for an agent, adding cases to an existing one, reviewing eval cases or scorers someone else wrote, choosing between a deterministic check and an LLM judge, deciding how many times to repeat a case, or reading a red run and working out whether the agent or the grader is wrong. Also use when a tool or prompt change "needs an eval" and it is not clear what…
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/elicitation.md`, `references/first-run.md` and `references/grading.md`).
It sits in AI & LLM Engineering, covering LLM evaluation, Building AI agents and Unit testing. It works with Model Context Protocol, LangChain, Langfuse and LangGraph. The repository describes itself as: Production-ready Next.js template for building AI agents with LangGraph.js. Features MCP integration for dynamic tool loading, human-in-the-loop tool approval, persistent… The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 40414f8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Eval Cases loads about 5.3k tokens when it runs, and up to ~24k if it reads all its reference files. Until then it costs about 192 tokens; SKILL.md has 3,366 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentailor/fullstack-langgraph-nextjs-agent at commit 40414f8, republished under its MIT licence (© agentailor). 3,366 words, ~5,284 tokens.
.claude/skills/agent-eval-cases/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.A case is one task you give the agent, plus the graders that decide whether what came back was acceptable.
That is the whole shape — and note what it is not. It is not an input paired with an expected output. There is no single correct answer string to compare against, which is why the right-hand side is a list of graders rather than a value. Many frameworks offer an expected / expected_output field; reaching for it by default is the most common way to write a suite that measures phrasing instead of behavior.
Three consequences shape everything below:
The hard part is not the format. It is knowing which handful of tasks are worth paying a model to run, repeatedly, forever.
Steps 1–3 are the ones that decide whether a suite is worth having. Do not skip to step 5.
A case needs something to run it on. Before writing any, find out what exists.
Look for an existing suite first. Search for an eval/evals/evaluation directory, a case or dataset type, *.eval.* files, or a framework dependency. If one exists, write cases in its idiom — its case type, its grader catalog, its repeat convention — and stop looking. See references/vocabulary.md.
If there is none, do not build one unprompted. Name the three options and let the builder choose:
| Option | When it fits |
|---|---|
| Grade by hand | A small suite. Run the agent, read the answer, decide if it is right. This is a real eval — a case, a run, and a human grader — and it is the correct starting point. |
| Build a minimal harness | Once run-count × case-count stops fitting in an afternoon. |
| Adopt a framework | When the plumbing (trace parsing, batching, CI reporting, dashboards) becomes the bottleneck rather than the point. |
Never scaffold a harness without explicit confirmation. Building one is a separate project with its own design decisions — what to capture, where it runs, how it resets. If the builder wants that, say so and get agreement first.
Hand-grading is not a lesser option to be talked out of. For a small agent it is the right answer, and it stays right for longer than people expect.
You cannot reason your way to a good case list from an empty page, and you must not try. A suite invented at a desk tests the failures someone was able to imagine. That is the wrong set.
Three sources, best first:
Ask; do not guess. If the builder cannot name a failure they have actually observed, say that plainly and offer a route to generate observations — a short hand-run probing session, or the domain-expert conversation — rather than inventing a suite that will look thorough and test nothing.
The question set, the triage, and what to do when there is nothing to go on are in references/elicitation.md.
This is the filter, and it is the step most likely to be skipped.
Observe a failure → push it to the cheapest layer that can catch it → let evals inherit only what will not fit.
A failure that a unit test could have caught is a failure you will pay to re-detect on every run, forever. Work down the layers:
tool-design territory.This filter assumes layer 2 exists. Check that it does. If the project has no tests over its tools, say so plainly — pushing a failure down to a layer that is not there means nothing catches it, and the case you were about to skip becomes the only guard. Two consequences worth stating to the builder:
Applying the fix and the case together is the honest sequence: a case has to be able to fail for the right reason before the defect underneath it is worth chasing.
Report what belongs at layers 1 and 2; do not go and build it. Those are changes to tool code and test suites, outside a request to write evals — and a payload fix changes the behavior every existing case was written against. Name the failure, say which layer should catch it and why, and let the builder decide whether to take it now, later, or not at all. Carry on writing cases for what is left.
What survives step 2 is a defect class — a way the agent can be wrong — not a feature. Grouping by feature produces one case per capability: expensive, slow, and mostly redundant with tests that already exist.
Defect classes look like this (examples of the shape, not a checklist to fill):
| Defect class | The failure | From |
|---|---|---|
| Tool mis-selection | Reaching for a listing tool to compute a total, when an aggregate was right there | analytics agent |
| Unnoticed truncation | Reporting a capped page as though it were the complete set | any paginated source |
| Unattributed currency | Repeating a dated source's snapshot as a fact about today | retrieval agent |
| Over-triggering | A caveat that was right once, now attached to every answer | retrieval agent |
| A guardrail or gate | A write reaching the store without the approval it required | human-in-the-loop agent |
| Prompt contracts | Instructions the prompt states in bold that nothing verifies | any agent |
| Silent bad data | A job that succeeds cleanly while writing the wrong thing | any agent that imports |
The right set for your agent is whatever step 2 left you, and it will not look like this one. A support agent groups around escalation and scope; a coding agent around destructive edits and test-passing-by-deletion. The axis is the defect, whatever the domain.
Keep the suite small. Five to ten cases is a normal, healthy first suite — start at the low end, since every case is one you pay for on every run. The anti-pattern is believing you need a big one before you start.
A case asserting "flag old sources as dated" invites a prompt fix that over-rotates into caveating everything — every answer hedged, the agent measurably worse, and the suite still green.
So each case that pushes a behavior gets a partner that bounds it. Here, that partner asserts a recent source is answered with no hedge. Where one case asserts a total is computed by aggregating, its partner asserts a plain listing still uses the listing tool.
Any rule that makes an agent do something needs a case proving it does not do it everywhere. An agent that qualifies every answer is worse than one that occasionally misses — and without the bounding case, that regression looks like total success.
First ask whether you can assert on what the agent did — a tool it called, a gate that fired, what the store holds afterwards. Those are structural facts with one spelling, and they often turn a judge-shaped target into a deterministic one: "asked before acting" needs a rubric, but the gate fired and the store is unchanged is two exact checks, and a stronger claim than any rubric would make.
For what is genuinely left, one question decides most of it:
Is the assertion target atomic?
A number, a tool name, a record count, a URL, an identifier has one spelling → a deterministic check is exact. A claim — "flagged the source as dated", "answered without hedging" — has combinatorially many correct spellings → it needs a judge.
Reach for a judge when the target genuinely has many correct spellings, and not before. Every judge is a second model call, a second thing that can be wrong, and a second thing to debug when a case goes red for no reason.
Having decided you need one, which judge is a second decision — judge calls are usually the dominant cost of a run, and a judge too small for the question produces false reds that look exactly like agent regressions.
Grader shapes, vacuous passes, rubric design, choosing and tiering the judge, and when to stop automating a case are in references/grading.md.
Non-determinism means one green run is weak evidence. A case can run n times and require k of those to pass. The two are set by different things, and conflating them is the usual mistake.
k follows grader stability. A deterministic grader returns the same verdict on the same capture, so any failure is real — set k = n. A judge is itself non-deterministic and has a small irreducible flip rate, so a lone misfire should not fail the case — set k = n - 1. A case mixing both still catches a real defect: a genuine regression fails its deterministic grader on every run.
n follows what a wrong answer costs. This is a budget decision, not a technical one, and it is where two agents legitimately diverge. An agent answering questions about published articles can afford to be wrong occasionally; an agent that moves money, sends mail, or writes to a system of record cannot. Start n small and raise it where the consequence justifies the spend — not everywhere.
Some cases get no tolerance at all. A guardrail, an approval gate, anything whose failure is silent or unrecoverable: every run must pass. Two out of three is not a pass there — it is a bug that happened to be outvoted.
| Policy | Shape | When |
|---|---|---|
| single | n=1, k=1 | Deterministic assertions on stable behavior; also the whole suite while iterating |
| majority | n=3, k=2 | Judge-graded cases, and behavior known to vary run to run |
| strict | n=3, k=3 | Deterministic graders you want repeated, guardrails, and outcomes that are silent or unrecoverable |
n is also a per-run mode, not only a per-case constant. While iterating on a case you expect to be red, confirming it red three times costs three times as much for no extra information. Run the suite at n=1 during the inner loop, then at full n before trusting the result. If the harness supports it, let cases derive n from a run-level policy rather than hardcoding it — a baked-in count silently defeats the fast mode.
Green under a reduced-
nrun is not a verdict. One run cannot distinguish a real pass from a lucky one. Re-run at fullnbefore gating on it, and record which mode produced a report — an iteration run that is later mistaken for a gate run is a hard mistake to catch.
Repeats are not noise reduction. The instinct is to run a flaky case a few times and take the majority so noise stops deciding the result. That is backwards — the flapping is the result. Run once and you learn "broken" or "fine" depending on which run you caught. Run several times and you learn the rate, which is what belongs in the bug report: "sometimes reports an empty result as zero" and "always does" are different bugs with different priorities.
Repeats multiply the bill, so spend them where instability is the point. A suite that gates merges wants them, since an unrepeated green cannot be told apart from a lucky one. A suite that exists to be read does not.
Writing the cases is not the end of the job. Most of a first suite's value shows up here.
Before the run:
Reading the run:
Report the run; do not act unilaterally on it. Say what each red means — which of the four suspects it points at, and what would have to change — then let the builder decide what gets fixed and in what order. As in step 2, prompt and tool changes sit outside a request to write evals, and a fix applied mid-suite changes the baseline every other case was written against.
The detail — including how to tell a legitimate grader fix from moving the goalposts — is in references/first-run.md.
These make a suite worse than having none, because each one produces confidence without information:
These make a suite cost more than it should, or measure the wrong thing:
Worth checking in a harness you did not write — report these rather than fixing them unasked:
© agentailor, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in .agents/skills/agent-eval-cases of agentailor/fullstack-langgraph-nextjs-agent.
Open the folder on GitHubat commit 40414f8
Agent Eval Cases next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Eval Cases this skillagentailor/fullstack-langgraph-nextjs-agent | 132 | — | ~5.3k | Automated safety check: Pass | MIT | |
| Dive Into LangGraphluochang212/dive-into-langgraph | 457 | — | ~837 | Automated safety check: Notes | Custom licence | |
| Agent Inspectrajudandigam/agent-inspect | 165 | — | ~424 | Automated safety check: Pass | MIT | |
| Chemgraphargonne-lcf/ChemGraph | 162 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Solana Agent Kitinternet-court/internet-court-skill | 6.4k | 3 repos | ~3.8k | Automated safety check: Notes | Apache-2.0 | |
| Magic ResumeMagic-Resume/Magic-Resume | 101 | — | ~663 | Automated safety check: Pass | MIT |
luochang212/dive-into-langgraph
A Chinese-language guide and reference for building agents with LangGraph 1.0, from a first ReAct agent through middleware, memory, MCP, RAG and web search.
rajudandigam/agent-inspect
Local evidence debugger and trajectory-test toolkit for TypeScript AI agents.
argonne-lcf/ChemGraph
Develop, test, and extend ChemGraph -- an agentic framework for automated molecular simulations using LLMs, LangGraph, ASE, and MCP servers
internet-court/internet-court-skill
Walks through building AI agents that run Solana operations such as token deploys, NFT minting, swaps and staking with SendAI's toolkit, in chat or fully autonomous mode.
Magic-Resume/Magic-Resume
How AI agents integrate with Magic Resume — read and safely edit a user's resumes through the native MCP server (@magic-resume/mcp).
pdovhomilja/nextcrm-app
Connect to NextCRM MCP server to manage CRM data — accounts, contacts, leads, opportunities, targets, products, contracts, activities, documents, target lists, enrichment, email accounts, campaigns…
agentailor/fullstack-langgraph-nextjs-agent
Comprehensive guide for designing, refining, and auditing system prompts for autonomous AI agents based on Anthropic's production practices.
agentailor/fullstack-langgraph-nextjs-agent
Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).
Categories
Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. Agent Eval Cases is an agent skill from agentailor/fullstack-langgraph-nextjs-agent. Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.
Agent Eval Cases fits situations like: writing a first eval suite for an agent; adding cases to an existing one; reviewing eval cases; scorers someone else wrote.
Run `npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a claude-code`. Or copy the skill folder (.agents/skills/agent-eval-cases in agentailor/fullstack-langgraph-nextjs-agent) into .claude/skills/agent-eval-cases in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a codex`. Or copy the skill folder (.agents/skills/agent-eval-cases in agentailor/fullstack-langgraph-nextjs-agent) into .agents/skills/agent-eval-cases in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill agent-eval-cases -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-eval-cases, .gemini/skills/agent-eval-cases, .github/skills/agent-eval-cases and .opencode/skills/agent-eval-cases in your project.
SKILL.md names no scripts, command-line tools or credentials: Agent Eval Cases is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agent Eval Cases is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agent Eval Cases: Dive Into LangGraph (luochang212/dive-into-langgraph, 457 stars), Agent Inspect (rajudandigam/agent-inspect, 165 stars), Chemgraph (argonne-lcf/ChemGraph, 162 stars) and Solana Agent Kit (internet-court/internet-court-skill, 6.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentailor (a GitHub organization) maintains it in agentailor/fullstack-langgraph-nextjs-agent, which has 132 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 4, 2026.
Source: agentailor/fullstack-langgraph-nextjs-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.