Agent skill

Palantir Incident Runbook

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Triage and stabilize Foundry API, pipeline, application, or Compute Module incidents with evidence-preserving rollback.

MITAuto-check passedDevOps & Cloud

Install Palantir Incident Runbook

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill palantir-incident-runbook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace palantir-incident-runbook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/palantir-incident-runbook .claude/skills/palantir-incident-runbook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
palantir-incident-runbook
GitHub stars
2.8k
Token cost
~1.4k tokens
SKILL.md length
613 words
Files
2 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Triage and stabilize Foundry API, pipeline, application, or Compute Module incidents with evidence-preserving rollback.

  • Works in 5 steps: Declare severity, impacted workflow,… → Classify the incident as API/auth,… → Preserve bounded evidence and freeze… → …
  • Degraded service
  • SKILL.md covers Overview, Prerequisites, Current Contract and Authentication, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Palantir Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Triage and stabilize Foundry API, pipeline, application, or Compute Module incidents with evidence-preserving rollback. Use when degraded service, failed builds, access failures, or release regressions require response. Trigger with "Foundry incident" or "Palantir outage".

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/official-docs.md`). Compatibility notes: Requires current Palantir Foundry documentation and approved access for any live resource, permission, data, build, application, or deployment change

It sits in DevOps & Cloud, covering Runbooks and postmortems and Incident response. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Degraded service
  • Access failures
  • Release regressions require response
  • With Foundry incident

Example prompts

  • “Foundry incident”
  • “Palantir outage”
  • “/palantir-incident-runbook”

Requirements

  • Compatibility (from SKILL.md): Requires current Palantir Foundry documentation and approved access for any live resource, permission, data, build, application, or deployment change
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, Write, Edit

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Declare severity, impacted workflow, current data-integrity risk, access-control risk, and the last known healthy version or build.
  2. Classify the incident as API/auth, transform/build, OSDK/application, Compute Module, data quality, permissions, or release regression.
  3. Preserve bounded evidence and freeze unrelated changes; restrict retries and writeback when integrity is uncertain.
  4. Choose the safest reversible mitigation: reduce concurrency, pause a schedule, disable a failing consumer, restore an artifact/product…
  5. Verify recovery through service health plus data correctness and permission checks, then publish a timeline and corrective actions.

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires current Palantir Foundry documentation and approved access for any live resource, permission, data, build, application, or deployment change

    From compatibility in the SKILL.md frontmatter.

Context cost

Palantir Incident Runbook loads about 1.4k tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 613 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 613 words, ~1,411 tokens.

Download SKILL.mdSave it as .claude/skills/palantir-incident-runbook/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
palantir-incident-runbook
description
Triage and stabilize Foundry API, pipeline, application, or Compute Module incidents with evidence-preserving rollback. Use when degraded service, failed builds, access failures, or release regressions require response. Trigger with "Foundry incident" or "Palantir outage".
allowed-tools
Read, Glob, Grep, Write, Edit
compatibility
Requires current Palantir Foundry documentation and approved access for any live resource, permission, data, build, application, or deployment change
version
2.0.0
argument-hint
[incident-or-resource-id]
model
inherit
effort
high
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
saas, palantir, foundry, incident-response, reliability

Palantir Foundry Incident Response

Overview

Protect data integrity and access controls while restoring service. Classify the failing Foundry surface, preserve request/build/release evidence, apply the least risky reversible mitigation, and validate recovery before closing.

Prerequisites

  • Name the incident commander, technical owner, data owner, communications owner, severity, affected resources, and start time.
  • Capture request IDs, build IDs, branch/commit, product or artifact version, module state, and recent approved changes.
  • Read references/official-docs.md and the target enrollment's operational procedures.
  • Confirm the rollback authority and protected-data handling rules before collecting logs.

Current Contract

  • API 429 or 503 can represent rate or concurrency limiting and should receive bounded backoff rather than an unbounded retry storm.
  • Transform build reports and metrics distinguish queue, CPU, memory, dependency, and data-related failures.
  • DevOps/Marketplace release management can retain prior versions and support controlled upgrades or rollback.
  • Logs may expose sensitive content and require explicit log access plus appropriate markings.

Authentication

Record grant type, principal, scope names, and application restrictions for API incidents, never token values. Do not rotate credentials until evidence distinguishes an authentication failure from permissions, restrictions, mandatory controls, throttling, or platform availability.

Instructions

  1. Declare severity, impacted workflow, current data-integrity risk, access-control risk, and the last known healthy version or build.

  2. Classify the incident as API/auth, transform/build, OSDK/application, Compute Module, data quality, permissions, or release regression.

  3. Preserve bounded evidence and freeze unrelated changes; restrict retries and writeback when integrity is uncertain.

  4. Choose the safest reversible mitigation: reduce concurrency, pause a schedule, disable a failing consumer, restore an artifact/product version, or isolate a bad output.

  5. Verify recovery through service health plus data correctness and permission checks, then publish a timeline and corrective actions.

Tool Discipline

  • Use Glob to locate candidate repositories, manifests, configurations, and evidence without widening scope.
  • Use Grep to find relevant identifiers, declarations, permissions, errors, and stale claims.
  • Use Read to inspect the smallest required files and authoritative evidence.
  • Use Write only for a new approved local draft, test, manifest, or evidence artifact.
  • Use Edit only for a bounded approved change whose rollback is known.
  • Do not use file tools as a substitute for authenticated Foundry operations or owner approval.
Show full SKILL.md (258 more words)Show less

Approval Boundaries

Only authorized owners may pause production schedules, change access, roll back products, alter module scaling, rotate credentials, or suppress outputs. Emergency authority must be recorded with scope and expiry.

Output

An incident record with severity, owners, timeline, evidence IDs, classification, customer/data impact, mitigation, approval, recovery validation, rollback outcome, and assigned corrective actions.

Error Handling

ConditionResponse
Evidence collection could expose protected dataUse identifiers and selected redacted excerpts; involve the data/security owner before accessing logs.
The last healthy version is unknownStop forward changes and reconstruct artifact/product/build history first.
Mitigation restores HTTP success but data is wrongKeep the incident open, isolate outputs, and validate lineage and transactions.
Repeated retries worsen impactCancel automation, reduce concurrency, and honor rate/concurrency guidance.

Examples

Example 1

Stabilize a backend integration returning 429 and 503 by halting duplicate workers, applying bounded backoff, checking current limits, and validating both request success and downstream object consistency.

Example 2

Roll back a DevOps product installation after an application release regression, verify the previous product version and OAuth/resource restrictions, and preserve the failed release evidence for follow-up.

Validation

  • Recovery is proven on the affected user or data workflow, not only a health endpoint.
  • Data integrity, access controls, and pending writeback are explicitly checked.
  • The exact mitigation and approval receipt are recorded.
  • Temporary emergency access or settings have been removed or assigned an expiry.
  • Corrective actions have owners and verification criteria.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/.curated/palantir-incident-runbook of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/official-docs.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Palantir Incident Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Palantir Incident Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Palantir Incident Runbook this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0
Activation Governance Chaos RolloutAli-Marandi/DataSense107—~1.9kAutomated safety check: PassMIT
Incident Response686f6c61/alfred-dev117—~1.1kAutomated safety check: PassMIT
Superset Incident Triagesuperset-sh/superset15k—~1kAutomated safety check: PassCustom licence
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT

Similar skills

  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Incident Response

    686f6c61/alfred-dev

    Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.

    117 GitHub stars~1.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Palantir Incident Runbook

What does Palantir Incident Runbook do?

Triage and stabilize Foundry API, pipeline, application, or Compute Module incidents with evidence-preserving rollback. Palantir Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Triage and stabilize Foundry API, pipeline, application, or Compute Module incidents with evidence-preserving rollback.

When should I use Palantir Incident Runbook?

Palantir Incident Runbook fits situations like: degraded service; access failures; release regressions require response; with Foundry incident.

How do I install Palantir Incident Runbook in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill palantir-incident-runbook -a claude-code`. Or copy the skill folder (skills/.curated/palantir-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/palantir-incident-runbook in your project. Claude Code loads it when a task matches its description.

How do I install Palantir Incident Runbook in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill palantir-incident-runbook -a codex`. Or copy the skill folder (skills/.curated/palantir-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/palantir-incident-runbook in your project. Codex loads it when a task matches its description.

Can I use Palantir Incident Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill palantir-incident-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/palantir-incident-runbook, .gemini/skills/palantir-incident-runbook, .github/skills/palantir-incident-runbook and .opencode/skills/palantir-incident-runbook in your project.

What does Palantir Incident Runbook need to run?

SKILL.md names no scripts, command-line tools or credentials: Palantir Incident Runbook is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Glob, Grep, Write, Edit. Compatibility (from SKILL.md): Requires current Palantir Foundry documentation and approved access for any live resource, permission, data, build, application, or deployment change.

Does Palantir Incident Runbook access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Palantir Incident Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Palantir Incident Runbook use?

Palantir Incident Runbook is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Palantir Incident Runbook use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 253 tokens, read only when the agent opens those files.

What are the alternatives to Palantir Incident Runbook?

Skills that share tags, products or a category with Palantir Incident Runbook: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Palantir Incident Runbook?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.