Agent Prompt Engineering
agentailor/fullstack-langgraph-nextjs-agent
Comprehensive guide for designing, refining, and auditing system prompts for autonomous AI agents based on Anthropic's production practices.
Tool skill — the first move in any LM-debugging session: dump the actual rendered prompt the framework sent to the model, before changing anything else.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-history-inspect --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-prompt-history-inspect .claude/skills/agentsop-prompt-history-inspect && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-prompt-history-inspect" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-history-inspect into .claude/skills/agentsop-prompt-history-inspect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-history-inspect", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-history-inspectType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-history-inspect --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-prompt-history-inspect .agents/skills/agentsop-prompt-history-inspect && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-prompt-history-inspect" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-history-inspect into .agents/skills/agentsop-prompt-history-inspect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-history-inspect", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-history-inspect --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-prompt-history-inspect .cursor/skills/agentsop-prompt-history-inspect && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-prompt-history-inspect" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-history-inspect into .cursor/skills/agentsop-prompt-history-inspect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-history-inspect", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-prompt-history-inspect--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-history-inspect --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-prompt-history-inspect .gemini/skills/agentsop-prompt-history-inspect && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-prompt-history-inspect" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-history-inspect into .gemini/skills/agentsop-prompt-history-inspect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-history-inspect", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-history-inspectInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-prompt-history-inspect .github/skills/agentsop-prompt-history-inspect && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-prompt-history-inspect" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-history-inspect into .github/skills/agentsop-prompt-history-inspect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-history-inspect", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-history-inspect --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-prompt-history-inspect .opencode/skills/agentsop-prompt-history-inspect && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-prompt-history-inspect" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-history-inspect into .opencode/skills/agentsop-prompt-history-inspect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-history-inspect", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-prompt-history-inspectTool skill — the first move in any LM-debugging session: dump the actual rendered prompt the framework sent to the model, before changing anything else.
Agentsop Prompt History Inspect is an agent skill from agentsope/SkillAlchemy. Tool skill — the first move in any LM-debugging session: dump the actual rendered prompt the framework sent to the model, before changing anything else. Activate when an LM call produced an unexpected output (wrong answer, schema violation, refusal, truncation, cost spike, latency spike, infinite loop, "model got dumber after upgrade"). The skill enforces a 30-second inspect step BEFORE any prompt edit, model swap, retry, or temperature tweak. Cross-framework cheat sheet: DSPy inspecthistory, LangGraph…
Its SKILL.md is about 8.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-mental-model.md`).
It sits in AI & LLM Engineering, covering Building AI agents and Prompt engineering. It works with OpenAI, CrewAI, LangChain and LangGraph. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.compython.langchain.comdspy.aidocs.crewai.comlangchain-ai.github.ioaider.chatdocs.langchain.comtil.simonwillison.netdocs.llamaindex.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Prompt History Inspect loads about 8.4k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 208 tokens; SKILL.md has 3,219 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 3,219 words, ~8,446 tokens.
.claude/skills/agentsop-prompt-history-inspect/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub."The prompt you wrote is not the prompt the model received." — Operating axiom for every framework that templates, injects few-shots, appends tool definitions, or wraps system messages.
Activate this skill the moment an LM call surprises you, BEFORE any other debug move.
| Trigger | Signal |
|---|---|
| Output wrong | "Why did it answer X?" / hallucinated fact / wrong format / refusal |
| Output truncated | mid-sentence cut, partial JSON, missing fields |
| Output empty / repeats | model returns "", repeats the same token, loops |
| Behaviour changed | "It worked yesterday" / "It worked on GPT-4o but not on Llama-3" |
| Cost / latency spike | tokens jumped 3× without code change → something got injected |
| Tool call wrong | wrong tool picked, args malformed, tool call missing |
| Schema validation failed | Pydantic / Outlines / guidance grammar refused output |
| Eval regression | metric dropped after upgrading framework version |
| Production bug | a user-facing thread produced a wrong answer — need to see what the LM saw |
Do NOT activate when:
The trigger is universal across the stack. Any framework that templates a prompt — DSPy, LangChain, LangGraph, CrewAI, LlamaIndex, Aider, Guidance, Outlines — has a layer between "what you wrote" and "what the model received." This skill is the first-line probe into that gap.
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ What you wrote │ ≠ │ What was rendered│ ≠ │ What the LM saw │
│ (template, │ │ (after few-shot │ │ (after provider │
│ signature, etc) │ │ injection, tool │ │ reformatting, │
│ │ │ defs appended, │ │ message squash, │
│ │ │ system msg, │ │ token truncation)
│ │ │ history, etc) │ │ │
└──────────────────┘ └──────────────────┘ └──────────────────┘
layer 1 layer 2 layer 3
(your code) (framework render) (provider transport)Three mental shifts the agent must internalize:
The rendered prompt is the ground truth, not your source code. Frameworks silently inject system messages, append tool JSON-schemas, deduplicate messages, summarise history, truncate context, reorder fields. The only trustworthy artefact is what was sent on the wire.
Inspect first, change nothing. Premature prompt edits hide the bug. If you "fix" the prompt before seeing the rendered output, you're optimising against a hallucination of the problem. The discipline is: dump → diff against expectation → identify which layer diverged → fix at that layer.
The inspect command is framework-specific but the SOP is universal. DSPy gives you inspect_history(n=1). LangChain gives you set_debug(True). Raw SDKs give you OPENAI_LOG=debug or HTTPX event hooks. You learn one cheat sheet (§7) once and the SOP applies everywhere.
The 30-second test. Before you do anything else, you should be able to print the exact final prompt text in <30 seconds. If you can't — you're not set up to debug LMs. Fix that first.
A four-step ritual. Do them in order. Do not skip ahead.
Look up the framework on the §7 cheat sheet. Run the inspect command. Get the actual text of:
For DSPy: dspy.inspect_history(n=1) [dspy.ai/api/utils/inspect_history/].
For LangChain: from langchain.globals import set_debug; set_debug(True) [python.langchain.com/api_reference/core/globals/langchain_core.globals.set_debug.html].
For raw SDKs: export OPENAI_LOG=debug or ANTHROPIC_LOG=debug [github.com/openai/openai-python, github.com/anthropics/anthropic-sdk-python].
Write down (mentally or in a scratch file): what did I expect the prompt to contain? Then compare to the dump. Look for:
None / empty / {var} literal).{% if %}).Now classify the bug into one of three layers (mental model §2):
max_tokens, restructure to avoid role coercion).Apply the minimum change at the identified layer. Re-run. Re-dump. Confirm the prompt now matches expectation. Then check whether the bug is fixed.
Critical anti-pattern: do not skip step 4's re-dump. Many "fixes" change Layer 1 when the real bug is Layer 2 — the prompt still looks broken on re-dump, even if the visible symptom changed.
Predict / ChainOfThought / ReAct returned wrong output, OR a compile run hung mid-trial.import dspy; dspy.inspect_history(n=1) — increase n to see the last few calls. For per-LM history: lm.inspect_history(n=3).inspect_history shows only LM calls, not retriever/tool calls — for those, wire mlflow.dspy.autolog() [github.com/stanfordnlp/dspy issue #784 for n-parameter quirks].from langchain.globals import set_debug, set_verbose
set_debug(True) # most verbose — full prompt + response + chain internals
# OR for less noise:
set_verbose(True) # prompt + response only, no chain internalshistory = list(graph.get_state_history(config)) # newest → oldest
for snap in history:
print(snap.config["configurable"]["checkpoint_id"], snap.values)
# Replay from a checkpoint:
graph.invoke(None, config={"configurable": {
"thread_id": "...", "checkpoint_id": "<chosen id>"}})
# Fork by modifying state:
graph.update_state(config, {"some_field": "new value"})Agent(..., verbose=True) # prints agent thoughts + tool I/O
Crew(..., verbose=True)
# Per-step structured capture:
def my_step(step_output):
print("STEP:", step_output) # AgentAction / AgentFinish / observation
Agent(..., step_callback=my_step)step_callback parameter), docs.crewai.com/en/observability/overview. CrewAI's internal log is thin — for production observability also wire mlflow.crewai.autolog() or Langtrace [docs.crewai.com/how-to/langtrace-observability]./diff # see exactly what Aider just changed
/tokens # see how big the context actually got
/ls # see which files are in chat vs read-onlyaider --verbose prints the full prompt sent to the model on each turn.--verbose) the rendered prompt.git diff HEAD~1 for verification — Aider auto-commits, so the diff is also in git history.client.chat.completions.create(...) (OpenAI) or client.messages.create(...) (Anthropic) directly, and the response is wrong. No framework templating layer to blame.# OpenAI:
export OPENAI_LOG=debug # also: OPENAI_LOG=info for less verbose
# Anthropic:
export ANTHROPIC_LOG=debuglogging module — the env var sets the logger level.debug level, the Authorization header (including the API key) is printed in plaintext [github.com/openai/openai-python issue #1196, issue #1082]. Never commit a debug-level log file. Strip keys before sharing.additional_drop_params or extra_body did), and OPENAI_LOG=debug is too noisy or formats badly.httpx.Client with event hooks to the SDK:import httpx, json
from openai import OpenAI
def log_request(req):
print("REQUEST:", req.method, req.url)
print(json.dumps(json.loads(req.content), indent=2))
client = OpenAI(http_client=httpx.Client(event_hooks={"request": [log_request]}))http_client=).langsmith, phoenix, langfuse) for the platform-specific UI. This skill enforces the first-move SOP; those skills provide the UI.tests/fixtures/prompt-bug-<id>.txt. Add a test that asserts the next render does not contain the bad pattern. (For LangChain, snapshot the prompt.format(**inputs) output; for DSPy, snapshot what inspect_history printed.)困境 (Dilemma): A RAG pipeline retrieves 10 passages and asks the LM to synthesize. Outputs miss obvious facts that are in the retrieved passages. User's first instinct: "the model is bad / the retriever is bad / let me re-rank."
约束 (Constraints):
MapReduceDocumentsChain wrapper.决策步骤 (Decision steps):
set_debug(True) from langchain.globals [python.langchain.com/api_reference/core/globals/langchain_core.globals.set_debug.html]. Re-run.stuff chain (all 10 passages in one prompt — fits in 128k) or write better chunk-summary prompts. Re-dump to confirm all 10 passages now appear.结果 (Outcome): Wrong layer would have been: a week of re-ranker tuning. Right layer: 10 minutes of template fix. Visible only via rendered-prompt dump.
可提取的操作 (Extractable operation): When a RAG pipeline misses obvious retrieved facts, dump the prompt and count how many retrieved chunks actually appear. The framework probably dropped some.
困境: A LangChain or CrewAI agent worked fine with 3 tools. After adding 4 more tools, accuracy dropped and latency tripled. User suspects the model "gets confused by more tools."
约束:
决策步骤:
tools=[...] array alone (OP-8).enum lists in params — describe them in natural language.结果: Without inspect, the user would "fix" by removing tools (losing capability) or switching models (expensive). Inspect reveals the tool-schema is the cost driver.
可提取的操作: More tools = silent prompt inflation. Always inspect tool-block token count before blaming the model.
困境: A MIPROv2-compiled program for GPT-4o hits 85% on dev. Re-pointed at Llama-3-8B, drops to 41%. User assumes Llama is just weaker.
约束:
compiled.json saved with demos + instructions tuned to GPT-4o.dspy.configure(lm=...) changed.决策步骤:
dspy.inspect_history(n=3) on Llama-3-8B [dspy.ai/api/utils/inspect_history/].compiled.json contain verbose, GPT-4o-style chains-of-thought (5–8 sentences per demo). Llama-3-8B copies the length but skips the reasoning structure — producing plausible-shaped but wrong outputs.dspy-sop]: recompile against Llama-3-8B. But without the inspect step, the user would not have known the demos were the bottleneck (vs. the instructions, or the signature).结果: Inspect reveals what changed in the rendered prompt; doctrine says what to do about it. Skipping inspect leads to "Llama is bad" — wrong root cause.
可提取的操作: Compiled-prompt artefacts are model-coupled. Inspect-dump on the new model is mandatory before declaring the model "weaker."
困境: A LangGraph customer-support agent gave a confidently wrong answer in production. User logs show only the final output, not intermediate state. No local repro.
约束:
决策步骤:
history = list(graph.get_state_history({"configurable": {"thread_id": "<prod-id>"}}))messages field in each StateSnapshot.values — that's the rendered prompt as seen by the LM [langchain-ai.github.io/langgraph/concepts/time-travel/].graph.invoke(None, config=...) to "replay" — that re-executes LM calls and incurs cost [time-travel docs caveat]. Reading state history is read-only and free.结果: Time-travel reads state without re-paying for LLM calls. The bug is visible in the inspected state, not in the final output alone.
可提取的操作: For production bugs, get_state_history is read-only and free; invoke(None, config=...) is replay and costs LM calls. Inspect first, replay only if necessary.
PromptTemplate.format(...) output is not the same as dumping what hit the wire — the framework adds messages, system instructions, tool schemas after that point. Prefer SDK-level (OPENAI_LOG=debug) or HTTPX-hook (OP-7) over template-render for production debugging.set_debug(True) / OPENAI_LOG=debug on in production. Both leak request bodies, and OPENAI_LOG=debug / ANTHROPIC_LOG=debug print the API key in plaintext [openai-python issue #1196]. Always scope to debug sessions; toggle off when done.graph.invoke(None, config=...) to debug. Re-executes LM calls and tools, paying real cost, possibly mutating real systems. Read state history; replay only when needed [time-travel docs].inspect_history as the only observability. DSPy's inspect_history shows LM calls only — not retrievers, not tools, not subgraphs. For multi-component pipelines, layer it with MLflow / LangSmith tracing [dspy.ai/tutorials/observability/].Authorization: headers before pasting into chat / issues / Slack.| Framework | First-move command | What it shows | Source |
|---|---|---|---|
| DSPy | dspy.inspect_history(n=1) | Last LM call: system / user / assistant / response | dspy.ai/api/utils/inspect_history/ |
| DSPy (per-LM) | lm.inspect_history(n=3) | Last N calls for a specific LM instance | dspy.ai/tutorials/observability/ |
| LangChain (max) | from langchain.globals import set_debug; set_debug(True) | Prompt + response + chain internals for every call | python.langchain.com/api_reference/core/globals/langchain_core.globals.set_debug.html |
| LangChain (lighter) | set_verbose(True) | Prompt + response only | python.langchain.com/api_reference/langchain/globals/langchain.globals.set_verbose.html |
| LangGraph | graph.get_state_history(config) | Per-step state including messages (rendered LM input) | langchain-ai.github.io/langgraph/concepts/time-travel/ |
| LangGraph (fork) | graph.update_state(config, {...}) then graph.invoke(None, config={"checkpoint_id": ...}) | Replay/fork from any past checkpoint (re-pays LM cost) | docs.langchain.com/oss/python/langgraph/use-time-travel |
| CrewAI | Agent(..., verbose=True) + Crew(..., verbose=True) | Agent thoughts, tool I/O, final answer | docs.crewai.com/en/concepts/agents |
| CrewAI (structured) | Agent(..., step_callback=fn) | Per-step AgentAction / observation captured in your callback | docs.crewai.com/en/concepts/agents, docs.crewai.com/en/observability/overview |
| LlamaIndex | Settings.callback_manager = CallbackManager([LlamaDebugHandler(...)]) then handler.get_llm_inputs_outputs() | All LLM inputs/outputs during a query | docs.llamaindex.ai (debugging guide) |
| Aider | /diff (last turn) + aider --verbose (full rendered prompt) | Per-turn diff and full prompt sent | aider.chat/docs/usage/commands.html |
| OpenAI SDK | export OPENAI_LOG=debug | Full HTTP request + response (incl. API key — strip!) | github.com/openai/openai-python README §Logging |
| Anthropic SDK | export ANTHROPIC_LOG=debug | Full HTTP request + response (incl. API key — strip!) | github.com/anthropics/anthropic-sdk-python README §Logging |
| Any SDK (clean) | Custom httpx.Client(event_hooks={"request": [log_fn]}) passed via http_client= | Structured JSON body of every request | til.simonwillison.net/httpx/openai-log-requests-responses |
| Production traces | Open the run in LangSmith / LangFuse / Phoenix / MLflow | Rendered prompt + response for a specific production thread | Per-platform skill |
LM call surprised you?
│
Yes ──► STOP. Do not edit the prompt. Do not retry. Do not swap models.
│
▼
What framework?
DSPy → dspy.inspect_history(n=1)
LangChain → set_debug(True)
LangGraph → graph.get_state_history(config)
CrewAI → step_callback + verbose=True
Aider → /diff AND aider --verbose
Raw OpenAI/Anth → OPENAI_LOG=debug / ANTHROPIC_LOG=debug
Clean dump → httpx event hook (OP-7)
Production bug → LangSmith / LangFuse / Phoenix trace
│
▼
Diff dump against expectation. Identify layer (1/2/3 — §3 Step 3).
│
▼
Fix at correct layer. Re-dump. Confirm prompt now matches.
│
▼
NOW check if the bug is fixed. If not, repeat from inspect.This skill is adjacent to but distinct from trace-UI skills:
langsmith, phoenix, langfuse, mlflow provide the UI for inspecting traces.prompt-history-inspect provides the SOP (inspect-first discipline) and the cross-framework command lookup.Use them together: this skill says when and why to inspect; the platform skills say where the UI lives. For local dev, the framework-native commands in §7 are usually enough.
inspect_history API: dspy.ai/api/utils/inspect_history/set_debug: python.langchain.com/api_reference/core/globals/langchain_core.globals.set_debug.htmlset_verbose: python.langchain.com/api_reference/langchain/globals/langchain.globals.set_verbose.htmlstep_callback, verbose): docs.crewai.com/en/concepts/agents© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/agentsop-prompt-history-inspect of agentsope/SkillAlchemy.
Open the folder on GitHubat commit d0f0355
Agentsop Prompt History Inspect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Prompt History Inspect this skillagentsope/SkillAlchemy | 436 | — | ~8.4k | Automated safety check: Pass | MIT | |
| Agent Prompt Engineeringagentailor/fullstack-langgraph-nextjs-agent | 132 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Mem0 Platform SDKmem0ai/mem0 | 67k | 1 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Add Example AgentGetBindu/Bindu | 10k | — | ~1.1k | Automated safety check: Notes | Custom licence | |
| Edgeone Makers MigrationTencentEdgeOne/edgeone-makers-tools | 1.9k | 1 repos | ~4.1k | Automated safety check: Pass | MIT | |
| Omnigent Framework Detectionomnigent-ai/omnigent | 11k | — | ~610 | Automated safety check: Pass | Apache-2.0 |
agentailor/fullstack-langgraph-nextjs-agent
Comprehensive guide for designing, refining, and auditing system prompts for autonomous AI agents based on Anthropic's production practices.
mem0ai/mem0
Adds persistent memory to AI apps with the Mem0 Python and TypeScript SDKs: store, search, update and delete user memories, with framework integrations.
GetBindu/Bindu
Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.
TencentEdgeOne/edgeone-makers-tools
Migrate existing AI agent projects (LangChain, LangGraph, OpenAI Agents SDK, Claude Agent SDK, CrewAI) to EdgeOne Makers platform conventions.
omnigent-ai/omnigent
Scans Python agent code for framework imports and recommends the matching Omnigent executor type, or says when the framework is not natively supported yet.
DataDog/dd-trace-js
A skill your agent uses when adding, debugging, or modifying LLMObs plugins for an LLM library in dd-trace-js.
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Categories
Tool skill — the first move in any LM-debugging session: dump the actual rendered prompt the framework sent to the model, before changing anything else. Agentsop Prompt History Inspect is an agent skill from agentsope/SkillAlchemy. Tool skill — the first move in any LM-debugging session: dump the actual rendered prompt the framework sent to the model, before changing anything else.
Agentsop Prompt History Inspect fits situations like: tasks that involve Building AI agents; tasks that involve Prompt engineering.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a claude-code`. Or copy the skill folder (skills/agentsop-prompt-history-inspect in agentsope/SkillAlchemy) into .claude/skills/agentsop-prompt-history-inspect in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a codex`. Or copy the skill folder (skills/agentsop-prompt-history-inspect in agentsope/SkillAlchemy) into .agents/skills/agentsop-prompt-history-inspect in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-history-inspect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-prompt-history-inspect, .gemini/skills/agentsop-prompt-history-inspect, .github/skills/agentsop-prompt-history-inspect and .opencode/skills/agentsop-prompt-history-inspect in your project.
Going by SKILL.md and its folder, Agentsop Prompt History Inspect needs the command-line tools its instructions call (git). Our summary lists: Python 3.
SKILL.md names 9 domains. As links in the text: github.com, python.langchain.com, dspy.ai, docs.crewai.com, langchain-ai.github.io, aider.chat, docs.langchain.com, til.simonwillison.net and docs.llamaindex.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Prompt History Inspect is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.4k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Prompt History Inspect: Agent Prompt Engineering (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Mem0 Platform SDK (mem0ai/mem0, 67k stars), Add Example Agent (GetBindu/Bindu, 10k stars) and Edgeone Makers Migration (TencentEdgeOne/edgeone-makers-tools, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 436 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.