Agent skill

Thought Based Reasoning

by NeoLabHQ in NeoLabHQ/context-engineering-kit

A skill your agent uses when tackling complex reasoning tasks requiring step-by-step logic, multi-step arithmetic, commonsense reasoning, symbolic manipulation, or problems where simple prompting…

GPL-3.0Auto-check passedAI & LLM Engineering

Install Thought Based Reasoning

skills CLI
$ npx skills add NeoLabHQ/context-engineering-kit --skill thought-based-reasoning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NeoLabHQ/context-engineering-kit thought-based-reasoning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/thought-based-reasoning .claude/skills/thought-based-reasoning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
thought-based-reasoning
GitHub stars
1.8k
Token cost
~5.5k tokens
SKILL.md length
1,478 words
Files
1
Skills in repo
57
Repo updated
First seen
Licence
GPL-3.0

At a glance

A skill your agent uses when tackling complex reasoning tasks requiring step-by-step logic, multi-step arithmetic, commonsense reasoning, symbolic manipulation, or problems where simple prompting…

  • Works in 12 steps: Chain-of-Thought (CoT) Prompting → Zero-shot Chain-of-Thought → Self-Consistency → …
  • Tackling complex reasoning tasks requiring step-by-step logic
  • SKILL.md covers Overview, Quick Reference, Core Techniques and Decision Matrix: Which…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Thought Based Reasoning is an agent skill from NeoLabHQ/context-engineering-kit. Use when tackling complex reasoning tasks requiring step-by-step logic, multi-step arithmetic, commonsense reasoning, symbolic manipulation, or problems where simple prompting fails - provides comprehensive guide to Chain-of-Thought and related prompting techniques (Zero-shot CoT, Self-Consistency, Tree of Thoughts, Least-to-Most, ReAct, PAL, Reflexion) with templates, decision matrices, and research-backed patterns

Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Prompt engineering. It works with React. The repository describes itself as: Hand-crafted Claude Code Skills focused on improving agent results quality. Compatible with OpenCode, Cursor, Antigravity, Gemini CLI, and others. Includes CodeRabbit open-source… The licence is GPL-3.0.

When your agent uses it

  • Tackling complex reasoning tasks requiring step-by-step logic
  • Multi-step arithmetic
  • Commonsense reasoning
  • Symbolic manipulation

Example prompts

  • “/thought-based-reasoning”

Requirements

  • Python 3

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Chain-of-Thought (CoT) Prompting
  2. Zero-shot Chain-of-Thought
  3. Self-Consistency
  4. Tree of Thoughts (ToT)
  5. Least-to-Most Prompting
  6. ReAct (Reasoning + Acting)
  7. PAL (Program-Aided Language Models)
  8. Auto-CoT
  9. Reflexion
  10. Start Simple
  11. Match Technique to Task
  12. Combine Techniques

What it can do on your machine

Read from SKILL.md and the folder at commit 23e2428. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Thought Based Reasoning loads about 5.5k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 1,478 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NeoLabHQ/context-engineering-kit at commit 23e2428, republished under its GPL-3.0 licence (© NeoLabHQ). 1,478 words, ~5,487 tokens.

Download SKILL.mdSave it as .claude/skills/thought-based-reasoning/SKILL.md (or your agent's skills folder).
name
thought-based-reasoning
description
Use when tackling complex reasoning tasks requiring step-by-step logic, multi-step arithmetic, commonsense reasoning, symbolic manipulation, or problems where simple prompting fails - provides comprehensive guide to Chain-of-Thought and related prompting techniques (Zero-shot CoT, Self-Consistency, Tree of Thoughts, Least-to-Most, ReAct, PAL, Reflexion) with templates, decision matrices, and research-backed patterns

Thought-Based Reasoning Techniques for LLMs

Overview

Chain-of-Thought (CoT) prompting and its variants encourage LLMs to generate intermediate reasoning steps before arriving at a final answer, significantly improving performance on complex reasoning tasks. These techniques transform how models approach problems by making implicit reasoning explicit.

Quick Reference

TechniqueWhen to UseComplexityAccuracy Gain
Zero-shot CoTQuick reasoning, no examples availableLow+20-60%
Few-shot CoTHave good examples, consistent format neededMedium+30-70%
Self-ConsistencyHigh-stakes decisions, need confidenceMedium+10-20% over CoT
Tree of ThoughtsComplex problems requiring explorationHigh+50-70% on hard tasks
Least-to-MostMulti-step problems with subproblemsMedium+30-80%
ReActTasks requiring external informationMedium+15-35%
PALMathematical/computational problemsMedium+10-15%
ReflexionIterative improvement, learning from errorsHigh+10-20%

Core Techniques

1. Chain-of-Thought (CoT) Prompting

Paper: "Chain of Thought Prompting Elicits Reasoning in Large Language Models" (Wei et al., 2022) Citations: 14,255+

When to Use
  • Multi-step arithmetic or math word problems
  • Commonsense reasoning requiring logical deduction
  • Symbolic reasoning tasks
  • When you have good exemplars showing reasoning
How It Works

Provide few-shot examples that include intermediate reasoning steps, not just question-answer pairs. The model learns to generate similar step-by-step reasoning.

Prompt Template
Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now?
A: Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11. The answer is 11.

Q: The cafeteria had 23 apples. If they used 20 to make lunch and bought 6 more, how many apples do they have?
A: The cafeteria had 23 apples originally. They used 20 to make lunch. So they had 23 - 20 = 3. They bought 6 more apples, so they have 3 + 6 = 9. The answer is 9.

Q: [YOUR QUESTION HERE]
A:
Strengths
  • Significant accuracy improvements on reasoning tasks
  • Interpretable intermediate steps
  • Works well with large models (>100B parameters)
Limitations
  • Requires crafting good exemplars
  • Less effective on smaller models
  • Can still make calculation errors

2. Zero-shot Chain-of-Thought

Paper: "Large Language Models are Zero-Shot Reasoners" (Kojima et al., 2022) Citations: 5,985+

When to Use
  • No exemplars available
  • Quick reasoning needed
  • General-purpose reasoning across task types
  • Prototyping before creating few-shot examples
How It Works

Simply append "Let's think step by step" (or similar phrase) to the prompt. This triggers the model to generate reasoning steps without any examples.

Prompt Template
Q: A juggler can juggle 16 balls. Half of the balls are golf balls, and half of the golf balls are blue. How many blue golf balls are there?

Let's think step by step.

Alternative trigger phrases:

  • "Let's work this out step by step to be sure we have the right answer."
  • "Let's break this down."
  • "Let's approach this systematically."
  • "First, let me understand the problem..."
Two-Stage Approach (More Robust)

Stage 1 - Reasoning Extraction:

Q: [QUESTION]
A: Let's think step by step.

Stage 2 - Answer Extraction:

[REASONING FROM STAGE 1]
Therefore, the answer is
Strengths
  • No exemplar crafting required
  • Generalizes across task types
  • Simple to implement
Limitations
  • Less effective than few-shot CoT
  • Can produce verbose or irrelevant reasoning
  • Sensitive to exact phrasing

3. Self-Consistency

Paper: "Self-Consistency Improves Chain of Thought Reasoning in Language Models" (Wang et al., 2022) Citations: 5,379+

When to Use
  • High-stakes decisions requiring confidence
  • Problems with multiple valid reasoning paths
  • When you need to reduce variance in outputs
  • Verification of reasoning correctness
How It Works

Sample multiple diverse reasoning paths, then select the most consistent answer via majority voting. The intuition: correct answers can be reached through multiple reasoning paths.

Prompt Template
[Use any CoT prompt - zero-shot or few-shot]

[Generate N samples with temperature > 0]

[Extract final answers from each sample]

[Return the most frequent answer (majority vote)]
Implementation Example
python
def self_consistency(prompt, n_samples=5, temperature=0.7):
    answers = []
    for _ in range(n_samples):
        response = llm.generate(prompt, temperature=temperature)
        answer = extract_answer(response)
        answers.append(answer)

    # Majority vote
    return Counter(answers).most_common(1)[0][0]
Strengths
  • Significant accuracy boost over single-path CoT
  • Provides confidence measure (agreement level)
  • Task-agnostic improvement
Limitations
  • Higher computational cost (N times more generations)
  • Requires extractable discrete answers
  • Diminishing returns beyond ~10-20 samples

4. Tree of Thoughts (ToT)

Paper: "Tree of Thoughts: Deliberate Problem Solving with Large Language Models" (Yao et al., 2023) Citations: 3,026+

When to Use
  • Complex problems requiring exploration/backtracking
  • Tasks where initial decisions are pivotal
  • Creative problem-solving (writing, puzzles)
  • When CoT alone achieves <50% accuracy
How It Works

Generalize CoT to a tree structure where each node is a "thought" (coherent language unit). Uses search algorithms (BFS/DFS) with self-evaluation to explore and select promising reasoning paths.

Prompt Template

Thought Generation:

Given the current state:
[STATE]

Generate 3-5 possible next steps to solve this problem.

State Evaluation:

Evaluate if the following partial solution is:
- "sure" (definitely leads to solution)
- "maybe" (could potentially work)
- "impossible" (cannot lead to solution)

Partial solution:
[THOUGHTS SO FAR]

BFS/DFS Search:

python
def tree_of_thoughts(problem, max_depth=3, beam_width=3):
    queue = [(problem, [])]  # (state, thought_path)

    while queue:
        state, path = queue.pop(0)

        if is_solved(state):
            return path

        # Generate candidate thoughts
        thoughts = generate_thoughts(state, k=5)

        # Evaluate and keep top-k
        evaluated = [(t, evaluate(state, t)) for t in thoughts]
        top_k = sorted(evaluated, key=lambda x: x[1])[:beam_width]

        for thought, score in top_k:
            if score != "impossible":
                new_state = apply_thought(state, thought)
                queue.append((new_state, path + [thought]))

    return None
Example: Game of 24
Problem: Use 4, 9, 10, 13 to get 24 (use +, -, *, / and each number once)

Thought 1: 13 - 9 = 4 (Now have: 4, 4, 10)
Evaluation: "maybe" - have two 4s and 10, could work

Thought 2: 10 - 4 = 6 (Now have: 4, 6, 13)
Evaluation: "maybe" - 4 * 6 = 24, need to use 13

Thought 3: 4 + 9 = 13 (Now have: 10, 13, 13)
Evaluation: "impossible" - no way to get 24 from these
Strengths
  • Dramatically improves performance on hard tasks (4% → 74% on Game of 24)
  • Enables backtracking and exploration
  • Self-evaluation catches errors early
Limitations
  • Significantly higher computational cost
  • Requires task-specific thought decomposition
  • Complex to implement

5. Least-to-Most Prompting

Paper: "Least-to-Most Prompting Enables Complex Reasoning in Large Language Models" (Zhou et al., 2022) Citations: 1,466+

When to Use
  • Problems harder than your exemplars
  • Compositional generalization tasks
  • Multi-step problems with clear subproblems
  • Symbol manipulation and SCAN-like tasks
How It Works

Two-stage process:

  1. Decomposition: Break complex problem into simpler subproblems
  2. Sequential Solving: Solve subproblems in order, using previous answers
Prompt Template

Stage 1: Decomposition

Q: Four years ago, Kody was only half as old as Mohamed. If Mohamed is currently twice as old as 30 years old, how old is Kody?

To solve "Four years ago, Kody was only half as old as Mohamed. If Mohamed is currently twice as old as 30 years old, how old is Kody?", we need to first solve:
- "If Mohamed is currently twice as old as 30 years old, how old is Mohamed?"
Then we need to solve:
- "Four years ago, Kody was only half as old as Mohamed. How old is Kody?"

Stage 2: Sequential Solving

Q: If Mohamed is currently twice as old as 30 years old, how old is Mohamed?
A: Mohamed is twice as old as 30, so Mohamed is 30 * 2 = 60 years old.

Q: Four years ago, Kody was only half as old as Mohamed. Mohamed is 60 years old. How old is Kody?
A: Four years ago, Mohamed was 60 - 4 = 56 years old. Kody was half as old, so Kody was 56 / 2 = 28 years old four years ago. Therefore, Kody is currently 28 + 4 = 32 years old.

The answer is 32.
Strengths
  • Excellent at generalizing to harder problems
  • Works well on compositional tasks
  • Explicit problem decomposition aids interpretability
Limitations
  • Requires two-stage prompting
  • Decomposition step can fail on novel structures
  • More complex setup than single-stage CoT

6. ReAct (Reasoning + Acting)

Paper: "ReAct: Synergizing Reasoning and Acting in Language Models" (Yao et al., 2022) Citations: 5,012+

When to Use
  • Tasks requiring external information (search, APIs)
  • Interactive decision-making environments
  • Multi-hop question answering
  • When pure reasoning leads to hallucination
How It Works

Interleave reasoning traces ("Thought") with actions ("Action") and observations ("Observation"). Reasoning helps plan actions; actions provide new information for reasoning.

Prompt Template
Question: What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into?

Thought 1: I need to search Colorado orogeny, find the area that the eastern sector extends into, then find the elevation range of that area.
Action 1: Search[Colorado orogeny]
Observation 1: The Colorado orogeny was an episode of mountain building (an orogeny) in Colorado and surrounding areas.

Thought 2: It does not mention the eastern sector. So I need to look up eastern sector.
Action 2: Lookup[eastern sector]
Observation 2: The eastern sector extends into the High Plains.

Thought 3: The eastern sector extends into the High Plains. I need to search High Plains and find its elevation range.
Action 3: Search[High Plains]
Observation 3: High Plains refers to one of two distinct land regions: High Plains (United States)...

Thought 4: I need to search High Plains (United States).
Action 4: Search[High Plains (United States)]
Observation 4: The High Plains are a subregion of the Great Plains... elevation of around 1,800 to 7,000 ft (550 to 2,130 m).

Thought 5: The elevation range is 1,800 to 7,000 ft.
Action 5: Finish[1,800 to 7,000 ft]
Action Types
  • Search[query] - Search for information
  • Lookup[keyword] - Look up keyword in current context
  • Finish[answer] - Return final answer
Strengths
  • Reduces hallucination by grounding in external knowledge
  • Interpretable action traces
  • Handles exceptions through adaptive reasoning
Limitations
  • Requires integration with external tools
  • More complex orchestration
  • Action space must be defined

7. PAL (Program-Aided Language Models)

Paper: "PAL: Program-aided Language Models" (Gao et al., 2022) Citations: 608+

When to Use
  • Mathematical/arithmetic reasoning
  • Problems requiring precise computation
  • Symbolic manipulation
  • When CoT makes calculation errors
How It Works

Generate code (typically Python) instead of natural language reasoning. Execute the code to get the answer. The LLM handles decomposition; the interpreter handles computation.

Prompt Template
Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now?

# solution in Python:
def solution():
    """Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now?"""
    tennis_balls_initial = 5
    bought_cans = 2
    tennis_balls_per_can = 3
    tennis_balls_bought = bought_cans * tennis_balls_per_can
    tennis_balls_total = tennis_balls_initial + tennis_balls_bought
    return tennis_balls_total

Q: The bakers at the Beverly Hills Bakery baked 200 loaves of bread on Monday morning. They sold 93 loaves in the morning and 39 loaves in the afternoon. A grocery store returned 6 unsold loaves. How many loaves of bread did they have left?

# solution in Python:
def solution():
    """The bakers baked 200 loaves. They sold 93 in morning, 39 in afternoon. A store returned 6. How many left?"""
    loaves_baked = 200
    loaves_sold_morning = 93
    loaves_sold_afternoon = 39
    loaves_returned = 6
    loaves_left = loaves_baked - loaves_sold_morning - loaves_sold_afternoon + loaves_returned
    return loaves_left
Strengths
  • Eliminates arithmetic errors
  • Clear variable naming aids interpretability
  • Leverages code execution for verification
Show full SKILL.md (598 more words)Show less
Limitations
  • Requires code interpreter
  • Not suitable for non-computational reasoning
  • Model must generate syntactically correct code

8. Auto-CoT

Paper: "Automatic Chain of Thought Prompting in Large Language Models" (Zhang et al., 2022) Citations: 838+

When to Use
  • No manually crafted exemplars available
  • Want to automate few-shot CoT setup
  • Scaling CoT to many tasks
  • When zero-shot CoT isn't sufficient
How It Works
  1. Cluster questions by diversity
  2. Use Zero-shot CoT to generate reasoning chains for representative questions
  3. Use these auto-generated chains as few-shot exemplars
Prompt Template

Step 1: Generate diverse demonstrations

python
# Cluster questions
clusters = cluster_questions(all_questions, k=8)

# For each cluster, pick representative and generate CoT
demonstrations = []
for cluster in clusters:
    question = select_representative(cluster)
    reasoning = zero_shot_cot(question)  # "Let's think step by step"
    demonstrations.append((question, reasoning))

Step 2: Use as few-shot exemplars

Q: [Demo question 1]
A: Let's think step by step. [Generated reasoning 1]

Q: [Demo question 2]
A: Let's think step by step. [Generated reasoning 2]

...

Q: [New question]
A: Let's think step by step.
Strengths
  • No manual exemplar creation
  • Diversity sampling improves robustness
  • Matches manual CoT performance
Limitations
  • Quality depends on zero-shot CoT quality
  • Clustering requires similarity metric
  • Some generated chains contain errors

9. Reflexion

Paper: "Reflexion: Language Agents with Verbal Reinforcement Learning" (Shinn et al., 2023) Citations: 2,179+

When to Use
  • Iterative improvement over multiple attempts
  • Learning from errors without fine-tuning
  • Complex coding or decision-making tasks
  • When single-pass reasoning is insufficient
How It Works

After task failure, the agent generates a verbal "reflection" analyzing what went wrong. This reflection is stored in memory and used in subsequent attempts to avoid repeating mistakes.

Prompt Template

Initial Attempt:

Task: [TASK DESCRIPTION]

Thought: [REASONING]
Action: [ACTION]
...
Result: [FAILURE/PARTIAL SUCCESS]

Reflection:

The previous attempt failed because:
1. [SPECIFIC ERROR ANALYSIS]
2. [WHAT SHOULD HAVE BEEN DONE]
3. [KEY INSIGHT FOR NEXT ATTEMPT]

Reflection: In the next attempt, I should...

Subsequent Attempt (with memory):

Task: [TASK DESCRIPTION]

Previous reflections:
- [REFLECTION 1]
- [REFLECTION 2]

Using these insights, I will now attempt the task again.

Thought: [IMPROVED REASONING]
Action: [BETTER ACTION]
Example: Code Generation
Task: Write a function to find the longest palindromic substring.

Attempt 1: [CODE WITH BUG]
Test Result: Failed on "babad" - expected "bab" or "aba", got "b"

Reflection: My solution only checked single characters. I need to:
1. Consider substrings of all lengths
2. Use expand-around-center technique for efficiency
3. Track both start position and maximum length

Attempt 2: [IMPROVED CODE USING REFLECTION]
Test Result: Passed all tests
Strengths
  • Learns from errors without weight updates
  • Achieves 91% on HumanEval (surpassing GPT-4's 80%)
  • Builds episodic memory of insights
Limitations
  • Requires multiple attempts
  • Memory management for long sessions
  • Quality of reflection affects improvement

Decision Matrix: Which Technique to Use

                           Need Examples?
                          /              \
                        No                Yes
                        |                  |
                Zero-shot CoT          Few-shot CoT
                        |                  |
                Need higher accuracy?  Need computation?
                /                \           |
              Yes               No          PAL
               |                |
    Self-Consistency    Done with CoT
               |
        Still not enough?
        /              \
      Yes              No
       |                |
  Problem decomposable?  Done
  /                    \
Yes                    No
 |                      |
Least-to-Most     Need exploration?
                  /              \
                Yes              No
                 |                |
          Tree of Thoughts   Need external info?
                             /              \
                           Yes              No
                            |                |
                          ReAct         Need iteration?
                                        /           \
                                      Yes           No
                                       |             |
                                   Reflexion      Use CoT

Best Practices

1. Start Simple

Begin with Zero-shot CoT ("Let's think step by step"), then progress to more complex techniques if needed.

2. Match Technique to Task
  • Math/Logic: CoT, PAL, Self-Consistency
  • Multi-hop QA: ReAct, Least-to-Most
  • Creative/Puzzles: Tree of Thoughts
  • Iterative Tasks: Reflexion
3. Combine Techniques

Techniques are often complementary:

  • ReAct + Self-Consistency for robust factual answers
  • ToT + PAL for complex computational exploration
  • Least-to-Most + Reflexion for hard multi-step problems
4. Prompt Engineering Tips
  • Use clear step markers ("Step 1:", "First,", etc.)
  • Include diverse exemplars covering edge cases
  • Format consistently across examples
  • Add verification steps ("Let me verify...")

Common Mistakes

MistakeWhy It's WrongFix
Using CoT for simple lookupsAdds unnecessary tokens and latencyReserve for multi-step reasoning
Too few samples in Self-ConsistencyMajority voting needs adequate samplesUse 5-10 samples minimum
Generic "think step by step" without checking outputModel may produce irrelevant reasoningValidate reasoning quality, not just presence
Mixing techniques without understanding trade-offsComputational cost without benefitUnderstand when each technique adds value
Using PAL without code interpreterCode generation is useless without executionEnsure execution environment available
Not testing exemplar quality in few-shot CoTPoor exemplars lead to poor reasoningValidate exemplars solve problems correctly
Applying Tree of Thoughts to linear problemsMassive overhead for no benefitUse ToT only when exploration needed

References

  1. Wei, J. et al. (2022). "Chain of Thought Prompting Elicits Reasoning in Large Language Models." arXiv:2201.11903

  2. Kojima, T. et al. (2022). "Large Language Models are Zero-Shot Reasoners." arXiv:2205.11916

  3. Wang, X. et al. (2022). "Self-Consistency Improves Chain of Thought Reasoning in Language Models." arXiv:2203.11171

  4. Yao, S. et al. (2023). "Tree of Thoughts: Deliberate Problem Solving with Large Language Models." arXiv:2305.10601

  5. Zhou, D. et al. (2022). "Least-to-Most Prompting Enables Complex Reasoning in Large Language Models." arXiv:2205.10625

  6. Yao, S. et al. (2022). "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv:2210.03629

  7. Gao, L. et al. (2022). "PAL: Program-aided Language Models." arXiv:2211.10435

  8. Zhang, Z. et al. (2022). "Automatic Chain of Thought Prompting in Large Language Models." arXiv:2210.03493

  9. Shinn, N. et al. (2023). "Reflexion: Language Agents with Verbal Reinforcement Learning." arXiv:2303.11366

© NeoLabHQ, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/thought-based-reasoning of NeoLabHQ/context-engineering-kit.

Open the folder on GitHubat commit 23e2428

Compare with similar skills

Thought Based Reasoning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Thought Based Reasoning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Thought Based Reasoning this skillNeoLabHQ/context-engineering-kit1.8k—~5.5kAutomated safety check: PassGPL-3.0
Canvas TemplatesPostHog/code179—~2.2kAutomated safety check: PassMIT
Building Agent Systemstelagod/code-abyss244—~691Automated safety check: PassMIT
Prompt Engineering Patternslamm-mit/scienceclaw246—~522Automated safety check: PassApache-2.0
Dspymagnus919/agent-skills115—~2kAutomated safety check: PassMIT
Prompt Improverseverity1/claude-code-prompt-improver1.9k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Canvas Templates

    PostHog/code

    Official

    How PostHog "canvas" dashboards work end-to-end — the two rendering tiers (json-render vs freeform React-in-iframe), the agent system prompts that steer each, and the RIGHT way to fetch PostHog data…

    179 GitHub stars~2.2k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Building Agent Systems

    telagod/code-abyss

    AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt…

    244 GitHub stars~691 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineering Patterns

    lamm-mit/scienceclaw

    Generate optimized LLM prompts using chain-of-thought, ReAct, and other scientific reasoning patterns

    246 GitHub stars~522 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dspy

    magnus919/agent-skills

    Optimize and build programmatic prompt systems with Stanford DSPy.

    115 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Prompt Improver

    severity1/claude-code-prompt-improver

    This skill enriches vague prompts with targeted research and clarification before execution.

    1.9k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineering Patterns

    ynulihao/AgentSkillOS

    Master advanced prompt engineering techniques to maximize LLM performance, reliability, and controllability in production.

    618 GitHub starsUsed in 14 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from NeoLabHQ/context-engineering-kit

All 57 skills in this repo
  • Git Notes

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when adding metadata to commits without changing history, tracking review status, test results, code quality annotations, or supplementing commit messages post-hoc - provides…

    1.8k GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Load PR Comments

    NeoLabHQ/context-engineering-kit

    A skill your agent uses to load open/unresolved PR review comments then aggregate them as tasks in .specs/comments/.md for parallel agents to fix.

    1.8k GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Prompt Engineering

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when you writing commands, hooks, skills for Agent, or prompts for sub agents or any other LLM interaction, including optimizing prompts, improving LLM outputs, or designing…

    1.8k GitHub stars~4.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Multi Agent Patterns

    NeoLabHQ/context-engineering-kit

    Design multi-agent architectures for complex tasks. An agent skill from NeoLabHQ/context-engineering-kit.

    1.8k GitHub starsUsed in 6 repos~6k tokens
    Auto-check passed
  • Review PR

    NeoLabHQ/context-engineering-kit

    Review an existing GitHub pull request and post inline review comments on its diff.

    1.8k GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Subagent Driven Development

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when executing implementation plans with independent tasks in the current session or facing 3+ independent issues that can be investigated without shared state or…

    1.8k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Thought Based Reasoning

What does Thought Based Reasoning do?

A skill your agent uses when tackling complex reasoning tasks requiring step-by-step logic, multi-step arithmetic, commonsense reasoning, symbolic manipulation, or problems where simple prompting…. Thought Based Reasoning is an agent skill from NeoLabHQ/context-engineering-kit.

When should I use Thought Based Reasoning?

Thought Based Reasoning fits situations like: tackling complex reasoning tasks requiring step-by-step logic; multi-step arithmetic; commonsense reasoning; symbolic manipulation.

How do I install Thought Based Reasoning in Claude Code?

Run `npx skills add NeoLabHQ/context-engineering-kit --skill thought-based-reasoning -a claude-code`. Or copy the skill folder (skills/thought-based-reasoning in NeoLabHQ/context-engineering-kit) into .claude/skills/thought-based-reasoning in your project. Claude Code loads it when a task matches its description.

How do I install Thought Based Reasoning in Codex?

Run `npx skills add NeoLabHQ/context-engineering-kit --skill thought-based-reasoning -a codex`. Or copy the skill folder (skills/thought-based-reasoning in NeoLabHQ/context-engineering-kit) into .agents/skills/thought-based-reasoning in your project. Codex loads it when a task matches its description.

Can I use Thought Based Reasoning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NeoLabHQ/context-engineering-kit --skill thought-based-reasoning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/thought-based-reasoning, .gemini/skills/thought-based-reasoning, .github/skills/thought-based-reasoning and .opencode/skills/thought-based-reasoning in your project.

What does Thought Based Reasoning need to run?

SKILL.md names no scripts, command-line tools or credentials: Thought Based Reasoning is instructions for the agent only. Our summary lists: Python 3.

Does Thought Based Reasoning access the network?

SKILL.md names 1 domain. As links in the text: arxiv.org. This is read from the text; nothing was executed.

Is Thought Based Reasoning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Thought Based Reasoning use?

Thought Based Reasoning is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Thought Based Reasoning use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Thought Based Reasoning?

Skills that share tags, products or a category with Thought Based Reasoning: Canvas Templates (PostHog/code, 179 stars), Building Agent Systems (telagod/code-abyss, 244 stars), Prompt Engineering Patterns (lamm-mit/scienceclaw, 246 stars) and Dspy (magnus919/agent-skills, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Thought Based Reasoning?

NeoLabHQ (a GitHub organization) maintains it in NeoLabHQ/context-engineering-kit, which has 1,750 GitHub stars. The repository holds 57 skills in this directory. The repository was last updated on August 26, 2026.

Source: NeoLabHQ/context-engineering-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.