Agent skill

Reasoning Serialization Tests

by tailcallhq in tailcallhq/forgecode

Checks that ReasoningConfig fields are serialized into the right provider-specific JSON for OpenRouter, Anthropic, GitHub Copilot and Codex requests.

Apache-2.0Auto-check passedTesting & QA

Install Reasoning Serialization Tests

skills CLI
$ npx skills add tailcallhq/forgecode --skill test-reasoning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tailcallhq/forgecode test-reasoning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tailcallhq/forgecode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.forge/skills/test-reasoning .claude/skills/test-reasoning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-reasoning
GitHub stars
7.6k
Token cost
~1k tokens
SKILL.md length
207 words
Files
2 (incl. scripts)
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Checks that ReasoningConfig fields are serialized into the right provider-specific JSON for OpenRouter, Anthropic, GitHub Copilot and Codex requests.

  • Verifying that reasoning settings reach the provider request correctly
  • SKILL.md covers Quick Start, Running a Single Test Manually, Test Coverage and References
  • Runs Shell scripts from its folder
  • After changing ReasoningConfig or a provider adapter

What it does

The bundled script scripts/test-reasoning.sh builds forge in debug mode, runs each provider and model combination, captures the outgoing HTTP request body through the FORGE_DEBUG_REQUESTS variable, and asserts that the expected JSON fields are present. To test one case by hand, you set that variable plus the provider and model IDs, run forge, and inspect .forge/forge.request.json.

A coverage table maps config fields to the JSON they should produce. For OpenRouter across several models, effort should appear as reasoning.effort, max_tokens as reasoning.max_tokens, enabled as reasoning.enabled, and exclude beside effort. For Anthropic, effort maps to output_config.effort. The table continues beyond what the excerpt shows.

When your agent uses it

  • Verifying that reasoning settings reach the provider request correctly
  • After changing ReasoningConfig or a provider adapter
  • Checking the reasoning JSON fields for a newly added model
  • Debugging a reasoning effort setting that seems to be ignored

Example prompts

  • “Run the reasoning serialization tests and report which providers fail.”
  • “Check that effort high on OpenRouter becomes reasoning.effort in the request body.”
  • “Test the Anthropic provider's effort mapping with FORGE_DEBUG_REQUESTS.”

Requirements

  • A forgecode checkout that builds in debug mode

What it can do on your machine

Read from SKILL.md and the folder at commit e11b93e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • developers.openai.com
    • platform.claude.com
    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reasoning Serialization Tests loads about 1k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 207 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tailcallhq/forgecode at commit e11b93e, republished under its Apache-2.0 licence (© tailcallhq). 207 words, ~1,037 tokens.

Download SKILL.mdSave it as .claude/skills/test-reasoning/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
test-reasoning
description
Validate that reasoning parameters are correctly serialized and sent to provider APIs. Use when the user asks to test reasoning serialization, run reasoning tests, verify reasoning config fields, or check that ReasoningConfig maps correctly to provider-specific JSON (OpenRouter, Anthropic, GitHub Copilot, Codex).

Test Reasoning Serialization

Validates that ReasoningConfig fields are correctly serialized into provider-specific JSON for OpenRouter, Anthropic, GitHub Copilot, and Codex.

Quick Start

Run all tests with the bundled script:

bash
./scripts/test-reasoning.sh

The script builds forge in debug mode, runs each provider/model combination, captures the outgoing HTTP request body via FORGE_DEBUG_REQUESTS, and asserts the correct JSON fields.

Running a Single Test Manually

bash
FORGE_DEBUG_REQUESTS="forge.request.json" \
FORGE_SESSION__PROVIDER_ID=<provider_id> \
FORGE_SESSION__MODEL_ID=<model_id> \
FORGE_REASONING__EFFORT=<effort> \
target/debug/forge -p "Hello!"

Then inspect .forge/forge.request.json for the expected fields.

Test Coverage

ProviderModelConfig fieldsExpected JSON field
open_routeropenai/o4-minieffort: none|minimal|low|medium|high|xhighreasoning.effort
open_routeropenai/o4-minimax_tokens: 4000reasoning.max_tokens
open_routeropenai/o4-minieffort: high + exclude: truereasoning.effort + .exclude
open_routeropenai/o4-minienabled: truereasoning.enabled
open_routeranthropic/claude-opus-4-5max_tokens: 4000reasoning.max_tokens
open_routermoonshotai/kimi-k2max_tokens: 4000reasoning.max_tokens
open_routermoonshotai/kimi-k2effort: highreasoning.effort
open_routerminimax/minimax-m2max_tokens: 4000reasoning.max_tokens
open_routerminimax/minimax-m2effort: highreasoning.effort
anthropicclaude-opus-4-6effort: low|medium|high|maxoutput_config.effort
anthropicclaude-3-7-sonnet-20250219enabled: true + max_tokens: 8000thinking.type + budget_tokens
github_copiloto4-minieffort: none|minimal|low|medium|high|xhighreasoning_effort (top-level)
codexgpt-5.1-codexeffort: none|minimal|low|medium|high|xhighreasoning.effort + .summary
codexgpt-5.1-codexeffort: medium + exclude: truereasoning.summary = "concise"
all providersone model eacheffort: invalidnon-zero exit, no request written

Tests for unconfigured providers are skipped automatically. Invalid-effort tests run regardless of credentials — the rejection happens at config parse time before any provider interaction.

References

© tailcallhq, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .forge/skills/test-reasoning of tailcallhq/forgecode.

  • SKILL.md
  • scripts/test-reasoning.sh

Open the folder on GitHubat commit e11b93e

Compare with similar skills

Reasoning Serialization Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reasoning Serialization Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reasoning Serialization Tests this skilltailcallhq/forgecode7.6k—~1kAutomated safety check: PassApache-2.0
PiDeck Usage Probe Helperayuayue/PiDeck1k—~1.4kAutomated safety check: PassMIT
LLM Modelsmajiayu000/claude-skill-registry6661 repos~992Automated safety check: PassMIT
Proxy Mode ReferenceMadAppGang/claude-code284—~1.3kAutomated safety check: PassMIT
Using Ccproxy Inspectorstarbaser/ccproxy350—~2.7kAutomated safety check: PassCustom licence
Using Ccproxy APIstarbaser/ccproxy350—~4kAutomated safety check: PassCustom licence

Similar skills

  • Helps show a model provider's usage, balance or quota in PiDeck: checks built-in support, points to the dialog templates, or writes a custom probe entry.

    1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Models

    majiayu000/claude-skill-registry

    Access Claude, Gemini, Kimi, GLM and 100+ LLMs via inference.sh CLI using OpenRouter.

    666 GitHub starsUsed in 1 repo~992 tokens
    AI & LLM EngineeringAuto-check passed
  • Proxy Mode Reference

    MadAppGang/claude-code

    Reference guide for using external AI models via claudish CLI.

    284 GitHub stars~1.3k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy Inspector

    starbaser/ccproxy

    Operates the ccproxy inspector MITM system for intercepting, inspecting, and transforming LLM API traffic.

    350 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy API

    starbaser/ccproxy

    Guides users through ccproxy as an OpenAI-compatible and Anthropic-compatible LLM API server with SDK integration, OAuth authentication, sentinel key substitution, model routing, and troubleshooting.

    350 GitHub stars~4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Keirouter Chat

    mydisha/keirouter

    Chat / code generation via KeiRouter using OpenAI /v1/chat/completions or Anthropic /v1/messages format with streaming + auto-fallback combos.

    147 GitHub stars~859 tokensUpdated 27 days ago
    AI & LLM EngineeringAuto-check passed

More from tailcallhq/forgecode

All 14 skills in this repo
  • Git Merge Conflict Resolver

    tailcallhq/forgecode

    Resolves Git merge conflicts with a plan-first workflow that keeps both sides' intent, regenerates lock files and backs up deleted-but-modified files.

    7.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Forge CLI Debug Workflow

    tailcallhq/forgecode

    Gives a systematic process for debugging the forge CLI: build in debug mode, check the latest help output, test with the non-interactive -p flag, and clone conversations before reproducing bugs.

    7.6k GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed
  • FIXME Resolver

    tailcallhq/forgecode

    Finds every FIXME comment in a codebase, groups related ones across files into one task, implements the work they describe and removes the comments once it is done.

    7.6k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Implementation Plan Creator

    tailcallhq/forgecode

    Writes a structured Markdown implementation plan with checkbox tasks, verification criteria and risks, then checks it with a validation script; no code changes.

    7.6k GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Release Notes Writer

    tailcallhq/forgecode

    Pulls a GitHub release and every linked pull request, then writes polished, factual release notes from the combined set of changes.

    7.6k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Skill Creation Guide

    tailcallhq/forgecode

    Guidance for creating and updating agent skills: what skills provide, keeping context lean, choosing how specific to be, and how SKILL.md and bundled resources are laid out.

    7.6k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Reasoning Serialization Tests

What does Reasoning Serialization Tests do?

Checks that ReasoningConfig fields are serialized into the right provider-specific JSON for OpenRouter, Anthropic, GitHub Copilot and Codex requests. sh builds forge in debug mode, runs each provider and model combination, captures the outgoing HTTP request body through the FORGE_DEBUG_REQUESTS variable, and asserts that the expected JSON fields are present.json.

When should I use Reasoning Serialization Tests?

Reasoning Serialization Tests fits situations like: verifying that reasoning settings reach the provider request correctly; after changing ReasoningConfig or a provider adapter; checking the reasoning JSON fields for a newly added model; debugging a reasoning effort setting that seems to be ignored.

How do I install Reasoning Serialization Tests in Claude Code?

Run `npx skills add tailcallhq/forgecode --skill test-reasoning -a claude-code`. Or copy the skill folder (.forge/skills/test-reasoning in tailcallhq/forgecode) into .claude/skills/test-reasoning in your project. Claude Code loads it when a task matches its description.

How do I install Reasoning Serialization Tests in Codex?

Run `npx skills add tailcallhq/forgecode --skill test-reasoning -a codex`. Or copy the skill folder (.forge/skills/test-reasoning in tailcallhq/forgecode) into .agents/skills/test-reasoning in your project. Codex loads it when a task matches its description.

Can I use Reasoning Serialization Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tailcallhq/forgecode --skill test-reasoning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-reasoning, .gemini/skills/test-reasoning, .github/skills/test-reasoning and .opencode/skills/test-reasoning in your project.

What does Reasoning Serialization Tests need to run?

Going by SKILL.md and its folder, Reasoning Serialization Tests needs a shell for the scripts in its folder. Our summary lists: A forgecode checkout that builds in debug mode.

Does Reasoning Serialization Tests access the network?

SKILL.md names 3 domains. As links in the text: developers.openai.com, platform.claude.com and openrouter.ai. This is read from the text; nothing was executed.

Is Reasoning Serialization Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Reasoning Serialization Tests use?

Reasoning Serialization Tests is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reasoning Serialization Tests use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Reasoning Serialization Tests?

Skills that share tags, products or a category with Reasoning Serialization Tests: PiDeck Usage Probe Helper (ayuayue/PiDeck, 1k stars), LLM Models (majiayu000/claude-skill-registry, 666 stars), Proxy Mode Reference (MadAppGang/claude-code, 284 stars) and Using Ccproxy Inspector (starbaser/ccproxy, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reasoning Serialization Tests?

tailcallhq (a GitHub organization) maintains it in tailcallhq/forgecode, which has 7,641 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.

Source: tailcallhq/forgecode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.