Official agent skill

Deploying Scalable Agents

by microsoft in microsoft/ai-agents-for-beginners

Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry.

OfficialMITAuto-check passedDevOps & Cloud

Install Deploying Scalable Agents

skills CLI
$ npx skills add microsoft/ai-agents-for-beginners --skill deploying-scalable-agents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/ai-agents-for-beginners deploying-scalable-agents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/ai-agents-for-beginners.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deploying-scalable-agents .claude/skills/deploying-scalable-agents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deploying-scalable-agents
GitHub stars
77k
Token cost
~1.5k tokens
SKILL.md length
600 words
Files
1
Skills in repo
123
Repo updated
First seen
Licence
MIT

At a glance

Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry.

  • Works in 3 steps: Client-hosted — the reasoning loop runs… → Hosted agent (Foundry Agent Service) —… → Agent workflow — multiple agents/tools…
  • : deploy an agent to production
  • SKILL.md covers Triggers, Core mental model, Deployment patterns (pick one,… and Lifecycle (the loop that ships…, plus 5 more sections
  • Calls az; reaches ai.azure.com

What it does

Deploying Scalable Agents is an agent skill from microsoft/ai-agents-for-beginners, published by the product's own GitHub organization. Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry. Covers deployment patterns (client-hosted, hosted agents, agent workflows), the agent lifecycle, model routing, response caching, evaluation gates, human-in-the-loop approval, observability with OpenTelemetry, cost optimisation, and smoke-testing deployed agents with the AI Smoke Test action. Based on Lesson 16 of AI Agents for Beginners. USE FOR: deploy an agent to production, scale an agent, Microsoft Foundry…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Observability, Caching and QA and bug reports. It works with Microsoft Azure and OpenTelemetry. The repository describes itself as: 18 Lessons to Get Started Building AI Agents. The licence is MIT.

When your agent uses it

  • : deploy an agent to production
  • Microsoft Foundry hosted agent
  • Foundry Agent Service
  • Response caching

Example prompts

  • “/deploying-scalable-agents”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Client-hosted — the reasoning loop runs in your process. Max control; you own scaling/state.
  2. Hosted agent (Foundry Agent Service) — Foundry hosts the loop, stores threads, enforces RBAC/content safety, shows the agent in the…
  3. Agent workflow — multiple agents/tools composed into a graph with branching, approval nodes, and durable checkpoints.

What it can do on your machine

Read from SKILL.md and the folder at commit 25b7985. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • az

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ai.azure.com

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deploying Scalable Agents loads about 1.5k tokens when it runs. Until then it costs about 253 tokens; SKILL.md has 600 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~253
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/ai-agents-for-beginners at commit 25b7985, republished under its MIT licence (© microsoft). 600 words, ~1,508 tokens.

Download SKILL.mdSave it as .claude/skills/deploying-scalable-agents/SKILL.md (or your agent's skills folder).
name
deploying-scalable-agents
description
Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry. Covers deployment patterns (client-hosted, hosted agents, agent workflows), the agent lifecycle, model routing, response caching, evaluation gates, human-in-the-loop approval, observability with OpenTelemetry, cost optimisation, and smoke-testing deployed agents with the AI Smoke Test action. Based on Lesson 16 of AI Agents for Beginners. USE FOR: deploy an agent to production, scale an agent, Microsoft Foundry hosted agent, Foundry Agent Service, model routing, response caching, evaluation gate, release gate, human approval workflow, agent observability, agent tracing, agent cost optimisation, smoke test a hosted agent, production customer support agent. DO NOT USE FOR: building your first agent (start with Lesson 01), running agents locally on-device (use local-ai-agents / Lesson 17), Azure infrastructure provisioning unrelated to agents, non-Foundry deployment targets.
license
MIT

Deploying Scalable Agents with Microsoft Foundry

Companion skill for Lesson 16 – Deploying Scalable Agents. Use it to help a learner move an agent from prototype to a scalable, observable production deployment. Ground every recommendation in the lesson content and the runnable notebook; do not invent Foundry APIs.

Triggers

Activate this skill when a learner wants to:

  • Deploy an agent to Microsoft Foundry as a hosted agent and make it versioned/observable.
  • Choose between client-hosted, hosted-agent, and agent-workflow deployment patterns.
  • Add model routing, response caching, or bounded concurrency to control latency and cost.
  • Add an evaluation gate so a bad agent version cannot ship.
  • Add a human-in-the-loop approval step for high-risk actions.
  • Instrument an agent with OpenTelemetry tracing for production observability.
  • Smoke-test a deployed agent as a fast post-deploy gate.

Core mental model

A production agent is mostly the operational skeleton around the model (~80%), not the model itself. Map every recommendation to one of these concerns:

ConcernPrototype → Production
Hostingnotebook → versioned hosted service
Identityyour az login → managed identity + scoped RBAC
Statein-memory → externalised thread/memory store
Failuretraceback → retries, fallbacks, alerts
Cost"a few cents" → tracked, routed, cached, budgeted
Qualityeyeballing → automated evaluation gate
Trustyou approve → policy + human-in-the-loop

Deployment patterns (pick one, or combine)

  1. Client-hosted — the reasoning loop runs in your process. Max control; you own scaling/state.
  2. Hosted agent (Foundry Agent Service) — Foundry hosts the loop, stores threads, enforces RBAC/content safety, shows the agent in the portal. Less control, far less operational surface.
  3. Agent workflow — multiple agents/tools composed into a graph with branching, approval nodes, and durable checkpoints.

Lifecycle (the loop that ships an agent)

create → version → evaluate (gate) → deploy hosted → observe online → collect failures → repeat. Offline evaluation is a gate, not an afterthought — a version does not ship unless it clears the threshold. Online observability feeds real failures back into the offline test set.

Scaling and cost levers (in priority order)

  1. Right-size the model — use the smallest model that passes the evaluation gate.
  2. Route by complexity — small/fast model for simple requests, large model for real reasoning (DIY classifier or Foundry Model Router).
  3. Cache — serve near-duplicate requests without a model call.
  4. Stateless design + bounded concurrency — externalise state; retry with backoff.
Show full SKILL.md (232 more words)Show less

Key patterns to reproduce

Point the learner at these from the notebook 16-python-agent-framework.ipynb:

  • Request handler: cache → route by complexity → trace span → run → cache.
  • Evaluation gate: score an offline test set; return pass_rate >= threshold and only deploy if true.
  • Human approval: @tool(approval_mode="always_require") for actions like large refunds.
  • Tracing: wrap each request in tracer.start_as_current_span(...) and set attributes like routed.model, customer.id.

Smoke-testing a deployed agent

After deploy, verify the endpoint actually answers (a green deploy can still be silent). Use the AI Smoke Test action via .github/workflows/smoke-test.yml with the catalog in tests/. The runner POSTs each prompt to POST {project_endpoint}/agents/{agent_name}/endpoint/protocols/openai/responses and asserts on the reply text. The identity needs the Azure AI User role at Foundry project scope; the token audience must be https://ai.azure.com/.

Layer the gates: smoke test (reachable/responding, every deploy) → offline evaluation (good enough to ship, before promotion) → online evaluation (how is it doing in the wild, continuous).

Enterprise controls

  • RBAC: give each hosted agent a managed identity with least privilege.
  • MCP in production: treat every MCP server as an untrusted boundary — pin the version, scope its identity, validate outputs, rate-limit, never expose secrets.

Guardrails for the assistant

  • Prefer the canonical FoundryChatClient(...) + provider.as_agent(...) pattern used across the course.
  • Do not promise live-Azure results you have not verified; recommend the smoke-test workflow to confirm a deployment.
  • Keep evaluation and cost advice tied together: evaluation sets the quality floor, routing/caching keep cost near that floor.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/deploying-scalable-agents of microsoft/ai-agents-for-beginners.

Open the folder on GitHubat commit 25b7985

Compare with similar skills

Deploying Scalable Agents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deploying Scalable Agents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deploying Scalable Agents this skillmicrosoft/ai-agents-for-beginners77k—~1.5kAutomated safety check: PassMIT
Aspiremicrosoft/aspire.dev1964 repos~1.1kAutomated safety check: PassMIT
Aspire MonitoringCommunityToolkit/Aspire629—~3.5kAutomated safety check: PassMIT
Temps Best Practicesgotempsh/temps822—~2.9kAutomated safety check: PassApache-2.0
Backdoor Deploymentmicrosoft/Docker-Provider173—~7kAutomated safety check: PassCustom licence
Observability Architecturemajiayu000/litellm-rs116—~1.3kAutomated safety check: PassMIT

Similar skills

  • Aspire

    microsoft/aspire.dev

    Official

    Orchestrates Aspire distributed applications using the Aspire CLI for running, debugging, and managing distributed apps.

    196 GitHub starsUsed in 4 repos~1.1k tokens
    DevOps & CloudAuto-check passed
  • Aspire Monitoring

    CommunityToolkit/Aspire

    ANALYSIS SKILL - Observe Aspire apps: logs, traces, metrics, resource state, telemetry export, browser telemetry, and the standalone dashboard.

    629 GitHub stars~3.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Temps Best Practices

    gotempsh/temps

    Best-practices reference for preparing and instrumenting applications on Temps.

    822 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Backdoor Deployment

    microsoft/Docker-Provider

    Official

    Validate a container image change via backdoor deployment. An agent skill from microsoft/Docker-Provider.

    173 GitHub stars~7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Observability Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

    116 GitHub stars~1.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Azure Monitor OpenTelemetry Exporter for Java. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 6 repos~2.2k tokens
    DevOps & CloudAuto-check passed

More from microsoft/ai-agents-for-beginners

All 123 skills in this repo
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    A skill your agent uses when the user asks to create, scaffold, or edit Jupyter notebooks (.ipynb) for experiments, explorations, or tutorials; prefer the bundled templates and run the helper script…

    77k GitHub starsUsed in 8 repos~1k tokens
    Auto-check passed
  • Azure Openai To Responses

    microsoft/ai-agents-for-beginners

    Official

    Migrate Python apps from Azure OpenAI Chat Completions to the Responses API.

    77k GitHub stars~6k tokensUpdated 18 days ago
    Auto-check: notes
  • Azure Openai To Responses

    microsoft/ai-agents-for-beginners

    Official

    Shift Python apps dem from Azure OpenAI Chat Completions go Responses API.

    77k GitHub stars~6k tokensUpdated 18 days ago
    Auto-check: notes
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    Kasuta, kui kasutaja palub luua, üles ehitada või redigeerida Jupyteri märkmikke (.ipynb) katsetuste, uurimiste või juhendite jaoks; eelista kaasasolevaid malle ja käivita abiskript newnotebook.py…

    77k GitHub stars~1.2k tokensUpdated 18 days ago
    Auto-check passed
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    Käytetään, kun käyttäjä pyytää luomaan, alustamaan tai muokkaamaan Jupyter-muistikirjoja (.ipynb) kokeita, tutkimuksia tai opetusohjelmia varten; käytä mieluummin mukana olevia mallipohjia ja…

    77k GitHub stars~1.3k tokensUpdated 18 days ago
    Auto-check passed
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    À utiliser lorsque l'utilisateur demande de créer, structurer ou modifier des notebooks Jupyter (.ipynb) pour des expériences, explorations ou tutoriels ; privilégiez les modèles fournis et exécutez…

    77k GitHub stars~1.4k tokensUpdated 18 days ago
    Auto-check passed

Questions about Deploying Scalable Agents

What does Deploying Scalable Agents do?

Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry. Deploying Scalable Agents is an agent skill from microsoft/ai-agents-for-beginners, published by the product's own GitHub organization. Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry.

When should I use Deploying Scalable Agents?

Deploying Scalable Agents fits situations like: : deploy an agent to production; microsoft Foundry hosted agent; foundry Agent Service; response caching.

How do I install Deploying Scalable Agents in Claude Code?

Run `npx skills add microsoft/ai-agents-for-beginners --skill deploying-scalable-agents -a claude-code`. Or copy the skill folder (.agents/skills/deploying-scalable-agents in microsoft/ai-agents-for-beginners) into .claude/skills/deploying-scalable-agents in your project. Claude Code loads it when a task matches its description.

How do I install Deploying Scalable Agents in Codex?

Run `npx skills add microsoft/ai-agents-for-beginners --skill deploying-scalable-agents -a codex`. Or copy the skill folder (.agents/skills/deploying-scalable-agents in microsoft/ai-agents-for-beginners) into .agents/skills/deploying-scalable-agents in your project. Codex loads it when a task matches its description.

Can I use Deploying Scalable Agents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/ai-agents-for-beginners --skill deploying-scalable-agents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deploying-scalable-agents, .gemini/skills/deploying-scalable-agents, .github/skills/deploying-scalable-agents and .opencode/skills/deploying-scalable-agents in your project.

What does Deploying Scalable Agents need to run?

Going by SKILL.md and its folder, Deploying Scalable Agents needs the command-line tools its instructions call (az).

Does Deploying Scalable Agents access the network?

SKILL.md names 2 domains. In commands or code: ai.azure.com; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Deploying Scalable Agents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deploying Scalable Agents use?

Deploying Scalable Agents is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deploying Scalable Agents use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deploying Scalable Agents?

Skills that share tags, products or a category with Deploying Scalable Agents: Aspire (microsoft/aspire.dev, 196 stars), Aspire Monitoring (CommunityToolkit/Aspire, 629 stars), Temps Best Practices (gotempsh/temps, 822 stars) and Backdoor Deployment (microsoft/Docker-Provider, 173 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deploying Scalable Agents?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/ai-agents-for-beginners, which has 76,567 GitHub stars. The repository holds 123 skills in this directory. The repository was last updated on September 19, 2026.

Source: microsoft/ai-agents-for-beginners on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.