Agent skill

Detecting Model Extraction Attacks

by mukul975 in mukul975/Anthropic-Cybersecurity-Skills

Detect MITRE ATLAS AML.T0024 attacks (model stealing, inversion, membership inference) performed via inference-API abuse, by monitoring per-principal query volume/distribution, rate-limiting and…

Apache-2.0Auto-check passedSecurity

Install Detecting Model Extraction Attacks

skills CLI
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-model-extraction-attacks -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills detecting-model-extraction-attacks --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/detecting-model-extraction-attacks .claude/skills/detecting-model-extraction-attacks && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
detecting-model-extraction-attacks
GitHub stars
34k
Token cost
~2.9k tokens
SKILL.md length
896 words
Files
5 (incl. scripts, references)
Skills in repo
639
Repo updated
First seen
Licence
Apache-2.0

At a glance

Detect MITRE ATLAS AML.T0024 attacks (model stealing, inversion, membership inference) performed via inference-API abuse, by monitoring per-principal query volume/distribution, rate-limiting and…

  • Works in 6 steps: Instrument the inference API for… → Detect extraction-like query patterns → Measure your model's extractability with… → …
  • Partner inference API needing cloning/inversion/membership-inference detection
  • SKILL.md covers Overview, When to Use, Prerequisites and Objectives, plus 5 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Detecting Model Extraction Attacks is an agent skill from mukul975/Anthropic-Cybersecurity-Skills. Detect MITRE ATLAS AML.T0024 attacks (model stealing, inversion, membership inference) performed via inference-API abuse, by monitoring per-principal query volume/distribution, rate-limiting and perturbing outputs, and red-teaming your model's extractability. Use for a public or partner inference API needing cloning/inversion/membership-inference detection, or a pre-deployment red-team exercise to measure extraction risk.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/api-reference.md`, `references/standards.md` and `scripts/agent.py`).

It sits in Security, covering Red teaming and adversary simulation and Rate limiting. The repository describes itself as: 817 structured cybersecurity skills for AI agents · Mapped to 6 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND, NIST AI RMF & MITRE F3 (Fight Fraud) · agentskills.io…. The licence is Apache-2.0.

When your agent uses it

  • Partner inference API needing cloning/inversion/membership-inference detection
  • A pre-deployment red-team exercise to measure extraction risk

Example prompts

  • “/detecting-model-extraction-attacks”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Instrument the inference API for detection signals
  2. Detect extraction-like query patterns
  3. Measure your model's extractability with ART (self red-team)
  4. Quantify training-data leakage with membership inference
  5. Apply and validate defenses
  6. Alert and respond

What it can do on your machine

Read from SKILL.md and the folder at commit 54a7988. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • atlas.mitre.org
    • github.com
    • adversarial-robustness-toolbox.readthedocs.io
    • nist.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Detecting Model Extraction Attacks loads about 2.9k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 896 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Anthropic-Cybersecurity-Skills at commit 54a7988, republished under its Apache-2.0 licence (© mukul975). 896 words, ~2,935 tokens.

Download SKILL.mdSave it as .claude/skills/detecting-model-extraction-attacks/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
detecting-model-extraction-attacks
description
Detect MITRE ATLAS AML.T0024 attacks (model stealing, inversion, membership inference) performed via inference-API abuse, by monitoring per-principal query volume/distribution, rate-limiting and perturbing outputs, and red-teaming your model's extractability. Use for a public or partner inference API needing cloning/inversion/membership-inference detection, or a pre-deployment red-team exercise to measure extraction risk.
domain
cybersecurity
subdomain
ai-security
tags
ai-security, model-extraction, membership-inference, model-inversion, inference-api, mitre-atlas, query-monitoring, mlsecops
version
1.0
author
mahipal
license
Apache-2.0
nist_ai_rmf
MEASURE-2.6
atlas_techniques
AML.T0024

Detecting Model Extraction Attacks

Authorized Use Only: The extraction, inversion, and membership-inference techniques described here are intended for defenders testing their own models and for red teams operating under written authorization. Querying a third-party model to clone it, reconstruct its training data, or infer membership without permission may violate terms of service, copyright, and privacy law.

Overview

Model extraction is the family of attacks in which an adversary abuses a model's inference API to steal value that the model owner intended to keep private. MITRE ATLAS catalogs these under AML.T0024 — Exfiltration via AI Inference API, in the Exfiltration tactic, with three sub-techniques:

  • AML.T0024.000 — Infer Training Data Membership (membership inference): the adversary determines whether a specific record was part of the training set, a privacy violation that can expose, for example, whether a patient's record trained a medical model.
  • AML.T0024.001 — Invert AI Model (model inversion): the adversary reconstructs representative training inputs (e.g., faces, text) by exploiting confidence scores returned by the API.
  • AML.T0024.002 — Extract ML Model (model stealing): the adversary repeatedly queries the victim model, collects (input, prediction) pairs, and trains a surrogate model offline that mimics the victim's decision boundary — avoiding the per-query cost of a Machine-Learning-as-a-Service offering and stealing the owner's intellectual property.

All three share a common signal: an attacker must send many queries, often crafted to probe the decision boundary (high-entropy, near-boundary, synthetic, or systematically grid-sampled inputs), and frequently requests full confidence vectors / logits rather than just the top label. Detection therefore centers on per-principal query monitoring, input-distribution analysis, and confidence-exposure controls, while defense centers on rate limiting, output perturbation, and reducing the information returned per query. This skill follows the MITRE ATLAS technique definition for AML.T0024 (https://atlas.mitre.org/techniques/AML.T0024) and the NIST AI RMF MEASURE function (MEASURE-2.6, security and resilience of the AI system).

When to Use

  • When you operate a model behind a public or partner inference API and need to detect cloning, inversion, or membership inference.
  • When performing a pre-deployment AI red-team exercise to measure how many queries are needed to extract your own model.
  • When validating that rate limiting, output perturbation, and confidence-suppression controls actually reduce extractability.
  • When investigating anomalous billing/usage spikes that may indicate surrogate-model harvesting.
  • When responding to a privacy incident where membership inference against a model is suspected.

Prerequisites

  • Python 3.9+ environment.
  • Access to inference-API access logs (per-API-key/per-principal query counts, timestamps, input features or hashes, returned confidence vectors).
  • For self-assessment red-teaming, install the Adversarial Robustness Toolbox (ART), the reference framework for extraction/inference attacks and defenses:
    bash
    pip install adversarial-robustness-toolbox scikit-learn numpy
  • Optional: access to the target model object (white/grey-box) or only its API (black-box).
  • Authorization to test the target model.

Objectives

  • Instrument the inference API to record per-principal query volume, input diversity, and confidence-exposure.
  • Build a detector that scores principals for extraction-like behavior (volume, near-boundary sampling, full-vector requests).
  • Run an ART-based extraction attack against your own model to measure fidelity vs. query budget.
  • Run a membership-inference attack to quantify training-data leakage.
  • Apply and validate defenses: rate limiting, label-only responses, confidence rounding/perturbation, and prediction poisoning.

MITRE ATT&CK Mapping

IDName (MITRE ATLAS)Tactic
AML.T0024Exfiltration via AI Inference APIExfiltration
AML.T0024.000Infer Training Data MembershipExfiltration
AML.T0024.001Invert AI ModelExfiltration
AML.T0024.002Extract ML ModelExfiltration

Workflow

Show full SKILL.md (368 more words)Show less
1. Instrument the inference API for detection signals

Capture the fields a detector needs. Per request, log the principal (API key / IP / account), timestamp, an input fingerprint, and whether the caller requested probabilities/logits.

python
import hashlib, json, time

def log_inference(principal, features, returned_probs):
    record = {
        "ts": time.time(),
        "principal": principal,
        # hash inputs so logs don't store raw sensitive data
        "input_hash": hashlib.sha256(json.dumps(features, sort_keys=True).encode()).hexdigest(),
        "wants_probs": returned_probs,
        "n_features": len(features),
    }
    with open("inference_audit.jsonl", "a") as f:
        f.write(json.dumps(record) + "\n")
2. Detect extraction-like query patterns

Score each principal on the three signals that distinguish extraction from normal use: high query volume in a window, high unique-input ratio (attackers rarely repeat), and a high rate of full-probability requests.

python
import collections, json

def score_principals(audit_path="inference_audit.jsonl", window_qps_threshold=100):
    by_principal = collections.defaultdict(lambda: {"q": 0, "uniq": set(), "probs": 0})
    for line in open(audit_path):
        r = json.loads(line)
        p = by_principal[r["principal"]]
        p["q"] += 1
        p["uniq"].add(r["input_hash"])
        p["probs"] += int(r["wants_probs"])
    findings = []
    for principal, p in by_principal.items():
        uniq_ratio = len(p["uniq"]) / max(p["q"], 1)
        prob_ratio = p["probs"] / max(p["q"], 1)
        suspicious = p["q"] > window_qps_threshold and uniq_ratio > 0.9 and prob_ratio > 0.8
        findings.append({"principal": principal, "queries": p["q"],
                         "unique_ratio": round(uniq_ratio, 3),
                         "prob_request_ratio": round(prob_ratio, 3),
                         "suspected_extraction": suspicious})
    return sorted(findings, key=lambda x: -x["queries"])
3. Measure your model's extractability with ART (self red-team)

Use ART's CopycatCNN (or KnockoffNets) to train a surrogate from black-box queries and report fidelity at a given query budget. Low query budget + high agreement = high risk.

python
import numpy as np
from art.estimators.classification import SklearnClassifier
from art.attacks.extraction import KnockoffNets
from sklearn.ensemble import RandomForestClassifier

# victim is your already-trained model wrapped for ART
victim = SklearnClassifier(model=trained_model)            # your production model
thief_model = RandomForestClassifier(n_estimators=100)
thief = SklearnClassifier(model=thief_model)

attack = KnockoffNets(classifier=victim, batch_size_fit=64,
                      batch_size_query=64, nb_epochs=10, nb_stolen=2000)
stolen = attack.extract(x=x_pool, thief_classifier=thief)   # 2000-query budget

agreement = np.mean(stolen.predict(x_test).argmax(1) == victim.predict(x_test).argmax(1))
print(f"Surrogate fidelity (agreement with victim): {agreement:.2%} at 2000 queries")
4. Quantify training-data leakage with membership inference

Run ART's black-box membership-inference attack. An accuracy meaningfully above 50% indicates the model leaks membership (AML.T0024.000).

python
from art.attacks.inference.membership_inference import MembershipInferenceBlackBox

mia = MembershipInferenceBlackBox(victim, attack_model_type="rf")
# fit the attack on a labeled split of known members / non-members
mia.fit(x_train[:500], y_train[:500], x_test[:500], y_test[:500])
member_pred = mia.infer(x_train[500:1000], y_train[500:1000])
nonmember_pred = mia.infer(x_test[500:1000], y_test[500:1000])
acc = (member_pred.mean() + (1 - nonmember_pred.mean())) / 2
print(f"Membership-inference accuracy: {acc:.2%} (0.50 = no leakage)")
5. Apply and validate defenses

Reduce the information returned and the query economics. Re-run steps 3 and 4 after each control to confirm extractability drops.

python
# (a) Label-only responses: never return full probability vectors to untrusted callers.
def respond(probs, trusted):
    return int(probs.argmax()) if not trusted else probs.tolist()

# (b) Confidence rounding / output perturbation (raises queries needed for inversion):
def perturb(probs, decimals=2, noise=0.01):
    p = np.round(probs, decimals) + np.random.normal(0, noise, probs.shape)
    p = np.clip(p, 0, None)
    return p / p.sum()

Defense in depth combines these with strict per-principal rate limiting, anomaly alerting from step 2, ART's ReverseSigmoid / prediction-poisoning postprocessor, and watermarking so an extracted surrogate remains attributable.

6. Alert and respond

Wire step-2 findings into your SIEM. On a confirmed extraction pattern: throttle or revoke the API key, switch the principal to label-only responses, preserve the audit log as evidence, and assess membership-inference exposure for any sensitive training data.

Tools and Resources

ResourceLink
MITRE ATLAS AML.T0024 — Exfiltration via AI Inference APIhttps://atlas.mitre.org/techniques/AML.T0024
Adversarial Robustness Toolbox (ART)https://github.com/Trusted-AI/adversarial-robustness-toolbox
ART extraction attacks (CopycatCNN, KnockoffNets)https://adversarial-robustness-toolbox.readthedocs.io/
MITRE ATLAS Matrixhttps://atlas.mitre.org/matrices/ATLAS
NIST AI RMF (MEASURE function)https://www.nist.gov/itl/ai-risk-management-framework

Detection Signal Reference

SignalNormal useExtraction behavior
Query volume per principalBounded, burstyVery high, sustained
Unique-input ratioRepeats common inputsNear-1.0 (rarely repeats)
Confidence-vector requestsMostly top labelDemands full probs/logits
Input distributionIn-distributionNear-boundary / synthetic / grid
Inter-query timingHuman-pacedAutomated, regular

Validation Criteria

  • Inference API logs per-principal query volume, input fingerprint, and confidence-exposure.
  • Detector scores principals and flags high-volume, high-unique-ratio, full-vector callers.
  • ART extraction attack run against own model; surrogate fidelity vs. query budget reported.
  • Membership-inference accuracy measured and compared against the 50% baseline.
  • Label-only / confidence-perturbation defenses applied and re-tested.
  • Per-principal rate limiting enforced and validated.
  • Alerts routed to SIEM with response playbook (throttle, revoke, preserve evidence).

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/detecting-model-extraction-attacks of mukul975/Anthropic-Cybersecurity-Skills.

  • SKILL.md
  • LICENSE
  • references/api-reference.md
  • references/standards.md
  • scripts/agent.py

Open the folder on GitHubat commit 54a7988

Compare with similar skills

Detecting Model Extraction Attacks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Detecting Model Extraction Attacks compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Detecting Model Extraction Attacks this skillmukul975/Anthropic-Cybersecurity-Skills34k—~2.9kAutomated safety check: PassApache-2.0
Identity Access Anomaly Reviewahmadvh/octochains375—~1.3kAutomated safety check: PassCustom licence
Convex Security Auditwaynesutton/builder-skills404—~2.6kAutomated safety check: PassApache-2.0
Operate Content Discoverycyberful/cyberful134—~1.5kAutomated safety check: PassAGPL-3.0
API Security Checklistrevfactory/harness-1001.3k—~1.7kAutomated safety check: PassApache-2.0
Secure Code GuardianJeffallan/claude-skills12k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Analyzes authentication and authorization events for failed-login clustering, privilege-escalation chains, credential-stuffing patterns, and MFA-bypass indicators.

    375 GitHub stars~1.3k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Convex Security Audit

    waynesutton/builder-skills

    Deep security review of a Convex app: authorization model, data access paths per table, HTTP action exposure, rate limiting, file storage access, scheduled function trust, and a written findings…

    404 GitHub stars~2.6k tokensUpdated 9 days ago
    SecurityAuto-check passed
  • Operate Content Discovery

    cyberful/cyberful

    Design and interpret advanced content discovery with ffuf and complementary web fuzzers.

    134 GitHub stars~1.5k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • API Security Checklist

    revfactory/harness-100

    Walks a backend-dev agent through OWASP API Top 10 checks, authentication and authorization patterns, and defense code during API design.

    1.3k GitHub stars~1.7k tokensUpdated 6 mo ago
    SecurityAuto-check passed
  • Secure Code Guardian

    Jeffallan/claude-skills

    Guides secure implementation of authentication, authorization, input validation and security headers, with password hashing, parameterized queries and OWASP Top 10 checks.

    12k GitHub stars~1.8k tokensUpdated 4 days ago
    SecurityAuto-check passed
  • Difesa Attacchi

    ccplugins/awesome-claude-code-plugins

    Aggiunge a un sito/app un agente di difesa che rileva e blocca richieste malevole (SQL injection, XSS, path traversal, brute force, bot) con rate limiting, blocklist IP e modalità lockdown che…

    967 GitHub stars~781 tokensUpdated 1 mo ago
    SecurityAuto-check passed

More from mukul975/Anthropic-Cybersecurity-Skills

All 639 skills in this repo
  • Campaign Attribution Evidence Analysis

    mukul975/Anthropic-Cybersecurity-Skills

    Weighs infrastructure, TTP, malware code and timing evidence with the Diamond Model and competing hypotheses to reach a confidence-rated attribution.

    34k GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Go Malware Analysis in Ghidra

    mukul975/Anthropic-Cybersecurity-Skills

    Walks through reverse engineering Go-compiled malware in Ghidra: parsing buildinfo and pclntab, recovering stripped function names and extracting dependencies.

    34k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • LNK and Jump List Forensics

    mukul975/Anthropic-Cybersecurity-Skills

    Guides forensic analysis of Windows LNK shortcut files and Jump Lists with LECmd, JLECmd and manual parsing to show file access and program execution.

    34k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Malware Persistence Analysis with Autoruns

    mukul975/Anthropic-Cybersecurity-Skills

    Hunts Windows malware persistence with Sysinternals Autoruns, covering run keys, services, scheduled tasks and drivers, with baseline comparison.

    34k GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • NTFS MFT Deleted File Recovery

    mukul975/Anthropic-Cybersecurity-Skills

    Guides a Windows forensic examination of the NTFS Master File Table to recover deleted-file evidence, build timelines and spot timestomping.

    34k GitHub stars~2.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Network Covert Channel Analysis

    mukul975/Anthropic-Cybersecurity-Skills

    Detects DNS tunneling, ICMP exfiltration and HTTP-based covert channels in packet captures and DNS logs when hunting for hidden command-and-control traffic.

    34k GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Detecting Model Extraction Attacks

What does Detecting Model Extraction Attacks do?

Detect MITRE ATLAS AML.T0024 attacks (model stealing, inversion, membership inference) performed via inference-API abuse, by monitoring per-principal query volume/distribution, rate-limiting and…. Detecting Model Extraction Attacks is an agent skill from mukul975/Anthropic-Cybersecurity-Skills.T0024 attacks (model stealing, inversion, membership inference) performed via inference-API abuse, by monitoring per-principal query volume/distribution, rate-limiting and perturbing outputs, and red-teaming your model's extractability.

When should I use Detecting Model Extraction Attacks?

Detecting Model Extraction Attacks fits situations like: partner inference API needing cloning/inversion/membership-inference detection; A pre-deployment red-team exercise to measure extraction risk.

How do I install Detecting Model Extraction Attacks in Claude Code?

Run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-model-extraction-attacks -a claude-code`. Or copy the skill folder (skills/detecting-model-extraction-attacks in mukul975/Anthropic-Cybersecurity-Skills) into .claude/skills/detecting-model-extraction-attacks in your project. Claude Code loads it when a task matches its description.

How do I install Detecting Model Extraction Attacks in Codex?

Run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-model-extraction-attacks -a codex`. Or copy the skill folder (skills/detecting-model-extraction-attacks in mukul975/Anthropic-Cybersecurity-Skills) into .agents/skills/detecting-model-extraction-attacks in your project. Codex loads it when a task matches its description.

Can I use Detecting Model Extraction Attacks in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-model-extraction-attacks -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/detecting-model-extraction-attacks, .gemini/skills/detecting-model-extraction-attacks, .github/skills/detecting-model-extraction-attacks and .opencode/skills/detecting-model-extraction-attacks in your project.

What does Detecting Model Extraction Attacks need to run?

Going by SKILL.md and its folder, Detecting Model Extraction Attacks needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Detecting Model Extraction Attacks access the network?

SKILL.md names 4 domains. As links in the text: atlas.mitre.org, github.com, adversarial-robustness-toolbox.readthedocs.io and nist.gov. This is read from the text; nothing was executed.

Is Detecting Model Extraction Attacks safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Detecting Model Extraction Attacks use?

Detecting Model Extraction Attacks is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Detecting Model Extraction Attacks use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 997 tokens, read only when the agent opens those files.

What are the alternatives to Detecting Model Extraction Attacks?

Skills that share tags, products or a category with Detecting Model Extraction Attacks: Identity Access Anomaly Review (ahmadvh/octochains, 375 stars), Convex Security Audit (waynesutton/builder-skills, 404 stars), Operate Content Discovery (cyberful/cyberful, 134 stars) and API Security Checklist (revfactory/harness-100, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Detecting Model Extraction Attacks?

mukul975 (a GitHub user) maintains it in mukul975/Anthropic-Cybersecurity-Skills, which has 33,870 GitHub stars. The repository holds 639 skills in this directory. The repository was last updated on August 31, 2026.

Source: mukul975/Anthropic-Cybersecurity-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.