Agent skill

Call Rlhf Self Reflection Scorer

by CALLE-AI in CALLE-AI/awesome-phone-call-agents

Offline experimental post-call feedback scorer using supplied ratings and text heuristics.

MITAuto-check passedAI & LLM Engineering

Install Call Rlhf Self Reflection Scorer

skills CLI
$ npx skills add CALLE-AI/awesome-phone-call-agents --skill call-rlhf-self-reflection-scorer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CALLE-AI/awesome-phone-call-agents call-rlhf-self-reflection-scorer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/CALLE-AI/awesome-phone-call-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/call-rlhf-self-reflection-scorer .claude/skills/call-rlhf-self-reflection-scorer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
call-rlhf-self-reflection-scorer
GitHub stars
107
Token cost
~727 tokens
SKILL.md length
343 words
Files
6 (incl. scripts, references)
Skills in repo
101
Repo updated
First seen
Licence
MIT

At a glance

Offline experimental post-call feedback scorer using supplied ratings and text heuristics.

  • Works in 4 steps: The skill receives the transcript and… → If the user gave a high score (>= 4),… → If the score is low or missing, the… → …
  • Demonstrate advisory review suggestions
  • SKILL.md covers Scientific Foundation, How it works, Decision Matrix and Expected Outcomes & Metrics, plus 1 more section
  • Runs Python scripts from its folder

What it does

Call Rlhf Self Reflection Scorer is an agent skill from CALLE-AI/awesome-phone-call-agents. Offline experimental post-call feedback scorer using supplied ratings and text heuristics. Use to demonstrate advisory review suggestions; no LLM inference, RLHF training, or memory integration is included.

Its SKILL.md is about 730 tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/examples.md`, `references/research-papers.md` and `references/safety.md`).

It sits in AI & LLM Engineering, covering Fine-tuning, Journaling and reflection and LLM inference and serving. The repository describes itself as: Portable phone-call Agent Skills, apps, examples, adapters, and scheduler recipes for AI agents. The licence is MIT.

When your agent uses it

  • Demonstrate advisory review suggestions
  • No LLM inference
  • Memory integration is included

Example prompts

  • “/call-rlhf-self-reflection-scorer”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. The skill receives the transcript and any explicit CSAT score given by the user (if applicable).
  2. If the user gave a high score (>= 4), the interaction is marked as successful.
  3. If the score is low or missing, the skill checks a small set of text patterns and returns a mocked critique; it cannot establish a root…
  4. It outputs an EvaluationResult containing the score, the identified critique, and a specific system prompt recommendation to fix the…

What it can do on your machine

Read from SKILL.md and the folder at commit 38d4118. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Call Rlhf Self Reflection Scorer loads about 727 tokens when it runs, and up to ~1.5k if it reads all its reference files. Until then it costs about 60 tokens; SKILL.md has 343 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~727
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from CALLE-AI/awesome-phone-call-agents at commit 38d4118, republished under its MIT licence (© CALLE-AI). 343 words, ~727 tokens.

Download SKILL.mdSave it as .claude/skills/call-rlhf-self-reflection-scorer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
call-rlhf-self-reflection-scorer
description
Offline experimental post-call feedback scorer using supplied ratings and text heuristics. Use to demonstrate advisory review suggestions; no LLM inference, RLHF training, or memory integration is included.
version
1.0.0

RLHF Self-Reflection Scorer

This skill demonstrates post-call QA with a small lexical scorer. It accepts a supplied transcript and optional rating, then returns predefined review suggestions. It does not collect user feedback, call an LLM, train a model, store RAG memories, or apply prompt changes. These are possible future host integrations, not delivered behavior.

Scientific Foundation

Paper / ConceptRelevance
LLM-as-a-JudgeUsing a strong LLM to evaluate the outputs of an agentic LLM correlates highly with human CSAT (Customer Satisfaction).
Self-Reflection (Reflexion)Agents that critique their own past transcripts and generate "verbal reinforcement" prompts perform significantly better on subsequent tasks.
RLHF (Reinforcement Learning from Human Feedback)Incorporating explicit user scores (if provided post-call) alongside automated critiques bridges the gap between simulated and real-world quality.

How it works

  1. The skill receives the transcript and any explicit CSAT score given by the user (if applicable).
  2. If the user gave a high score (>= 4), the interaction is marked as successful.
  3. If the score is low or missing, the skill checks a small set of text patterns and returns a mocked critique; it cannot establish a root cause.
  4. It outputs an EvaluationResult containing the score, the identified critique, and a specific system prompt recommendation to fix the behavior.

Decision Matrix

Explicit ScoreTranscript SentimentOutcomeAction
>= 4AnyEXPLICIT_USER (High)Maintain current strategy
< 4AnySELF_CRITIQUEReturn a heuristic critique; not an explanation of the user's rating
NoneSmoothSELF_CRITIQUE (High)Baseline evaluation
NoneFriction detectedSELF_CRITIQUE (Low)Flag friction point and generate patch

Expected Outcomes & Metrics

These are unvalidated targets for a possible future evaluator, not measurements of this mock scorer.

MetricTargetNotes
Critique Relevance> 90%The LLM-generated critique should match human QA audits.
Recommendation Actionability> 85%Recommendations must be directly usable as system prompt instructions.

Limitations & Known Constraints

  • Self-Correction Loop: This skill only generates the critique. A separate meta-agent is required to actually update the core agent's prompt based on these recommendations.
  • Cost: The local helper has no LLM dependency. A future LLM integration would have separate costs and evaluation needs.

© CALLE-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/call-rlhf-self-reflection-scorer of CALLE-AI/awesome-phone-call-agents.

  • SKILL.md
  • references/examples.md
  • references/research-papers.md
  • references/safety.md
  • scripts/self_scorer.py
  • scripts/test_self_scorer.py

Open the folder on GitHubat commit 38d4118

Compare with similar skills

Call Rlhf Self Reflection Scorer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Call Rlhf Self Reflection Scorer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Call Rlhf Self Reflection Scorer this skillCALLE-AI/awesome-phone-call-agents107—~727Automated safety check: PassMIT
Nemotron Super3NVIDIA-NeMo/Nemotron2.1k—~2.4kAutomated safety check: PassApache-2.0
GptqOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Openrlhf TrainingOrchestra-Research/AI-Research-SKILLs13k2 repos~2.1kAutomated safety check: NotesMIT
bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs13k2 repos~2.5kAutomated safety check: PassMIT
Nemotron UltraNVIDIA-NeMo/Nemotron2.1k—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Nemotron Super3

    NVIDIA-NeMo/Nemotron

    Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.

    2.1k GitHub stars~2.4k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Gptq

    Orchestra-Research/AI-Research-SKILLs

    Post-training 4-bit quantization for LLMs with minimal accuracy loss.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Openrlhf Training

    Orchestra-Research/AI-Research-SKILLs

    High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs.

    13k GitHub starsUsed in 2 repos~2.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • bitsandbytes Model Quantization

    Orchestra-Research/AI-Research-SKILLs

    Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

    13k GitHub starsUsed in 2 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Nemotron Ultra

    NVIDIA-NeMo/Nemotron

    Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference.

    2.1k GitHub stars~1.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub stars~1.7k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed

More from CALLE-AI/awesome-phone-call-agents

All 101 skills in this repo
  • Accessible Outing Verifier

    CALLE-AI/awesome-phone-call-agents

    Demonstrates advisory accessibility-planning checks with offline fixtures and a proposed bounded CALL-E workflow; use for exploring unknown or qualified venue claims without making calls.

    107 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Ground Truth Gate

    CALLE-AI/awesome-phone-call-agents

    A skill your agent uses when an agent holds some evidence for a physical-world claim but the evidence is broader, narrower, or older than the exact question asked, and it must first decide whether a…

    107 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Is It Accessible

    CALLE-AI/awesome-phone-call-agents

    Call a venue and ask the accessibility questions that matter to one specific person — step-free entry, hearing loop, guide dogs, quiet hours, changing places — then return a per-need verdict backed…

    107 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Landmark Navigation Assist

    CALLE-AI/awesome-phone-call-agents

    Turns a pre-written, building-level location config into a CALL-E outbound phone-call task that guides a delivery driver through the last few hundred metres to a specific building using landmarks…

    107 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Research Gap Call Verifier

    CALLE-AI/awesome-phone-call-agents

    Turn cited business research into a bounded, approval-gated phone-call plan that asks only unresolved factual questions, then reconcile CALL-E-compatible results without treating voicemail, refusal…

    107 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Structured Outcome Followup Call

    CALLE-AI/awesome-phone-call-agents

    Place a goal-driven CALL-E call that collects specific structured answers, score those answers against a deterministic rubric you supply, and conditionally trigger a follow-up action — all runnable…

    107 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Call Rlhf Self Reflection Scorer

What does Call Rlhf Self Reflection Scorer do?

Offline experimental post-call feedback scorer using supplied ratings and text heuristics. Call Rlhf Self Reflection Scorer is an agent skill from CALLE-AI/awesome-phone-call-agents. Offline experimental post-call feedback scorer using supplied ratings and text heuristics.

When should I use Call Rlhf Self Reflection Scorer?

Call Rlhf Self Reflection Scorer fits situations like: demonstrate advisory review suggestions; no LLM inference; memory integration is included.

How do I install Call Rlhf Self Reflection Scorer in Claude Code?

Run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-rlhf-self-reflection-scorer -a claude-code`. Or copy the skill folder (skills/call-rlhf-self-reflection-scorer in CALLE-AI/awesome-phone-call-agents) into .claude/skills/call-rlhf-self-reflection-scorer in your project. Claude Code loads it when a task matches its description.

How do I install Call Rlhf Self Reflection Scorer in Codex?

Run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-rlhf-self-reflection-scorer -a codex`. Or copy the skill folder (skills/call-rlhf-self-reflection-scorer in CALLE-AI/awesome-phone-call-agents) into .agents/skills/call-rlhf-self-reflection-scorer in your project. Codex loads it when a task matches its description.

Can I use Call Rlhf Self Reflection Scorer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-rlhf-self-reflection-scorer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/call-rlhf-self-reflection-scorer, .gemini/skills/call-rlhf-self-reflection-scorer, .github/skills/call-rlhf-self-reflection-scorer and .opencode/skills/call-rlhf-self-reflection-scorer in your project.

What does Call Rlhf Self Reflection Scorer need to run?

Going by SKILL.md and its folder, Call Rlhf Self Reflection Scorer needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Call Rlhf Self Reflection Scorer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Call Rlhf Self Reflection Scorer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Call Rlhf Self Reflection Scorer use?

Call Rlhf Self Reflection Scorer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Call Rlhf Self Reflection Scorer use?

About 727 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 789 tokens, read only when the agent opens those files.

What are the alternatives to Call Rlhf Self Reflection Scorer?

Skills that share tags, products or a category with Call Rlhf Self Reflection Scorer: Nemotron Super3 (NVIDIA-NeMo/Nemotron, 2.1k stars), Gptq (Orchestra-Research/AI-Research-SKILLs, 13k stars), Openrlhf Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and bitsandbytes Model Quantization (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Call Rlhf Self Reflection Scorer?

CALLE-AI (a GitHub organization) maintains it in CALLE-AI/awesome-phone-call-agents, which has 107 GitHub stars. The repository holds 101 skills in this directory. The repository was last updated on October 10, 2026.

Source: CALLE-AI/awesome-phone-call-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.