Official agent skill

Deploying Scalable Agents

by microsoft in microsoft/ai-agents-for-beginners

Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry.

OfficialMITAuto-check passedDevOps & Cloud

Install Deploying Scalable Agents

skills CLI
$ npx skills add microsoft/ai-agents-for-beginners --skill deploying-scalable-agents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/ai-agents-for-beginners deploying-scalable-agents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/ai-agents-for-beginners.git skills-src && mkdir -p .claude/skills && cp -r skills-src/translations/pcm/.agents/skills/deploying-scalable-agents .claude/skills/deploying-scalable-agents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deploying-scalable-agents
GitHub stars
77k
Token cost
~1.6k tokens
SKILL.md length
660 words
Files
1
Skills in repo
123
Repo updated
First seen
Licence
MIT

At a glance

Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry.

  • Works in 3 steps: Client-hosted — di reasoning loop dey… → Hosted agent (Foundry Agent Service) —… → Agent workflow — multiple agents/tools…
  • : deploy one agent go production
  • SKILL.md covers Triggers, Core mental model, Deployment patterns (pick one,… and Lifecycle (di loop wey dey…, plus 5 more sections
  • Calls az; reaches ai.azure.com

What it does

Deploying Scalable Agents is an agent skill from microsoft/ai-agents-for-beginners, published by the product's own GitHub organization. Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry. E cover deployment patterns (client-hosted, hosted agents, agent workflows), the agent lifecycle, model routing, response caching, evaluation gates, human-in-the-loop approval, observability with OpenTelemetry, cost optimisation, and smoke-testing deployed agents with the AI Smoke Test action. Based on Lesson 16 of AI Agents for Beginners. USE FOR: deploy one agent go production, scale one agent, Microsoft…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Observability, Caching and QA and bug reports. It works with Microsoft Azure and OpenTelemetry. The repository describes itself as: 18 Lessons to Get Started Building AI Agents. The licence is MIT.

When your agent uses it

  • : deploy one agent go production
  • Scale one agent
  • Microsoft Foundry hosted agent
  • Foundry Agent Service

Example prompts

  • “/deploying-scalable-agents”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Client-hosted — di reasoning loop dey run inside your process. Max control; you dey own scaling/state.
  2. Hosted agent (Foundry Agent Service) — Foundry dey host the loop, dey store threads, dey enforce RBAC/content safety, dey show di agent…
  3. Agent workflow — multiple agents/tools join body as graph wit branching, approval nodes, and durable checkpoints.

What it can do on your machine

Read from SKILL.md and the folder at commit 25b7985. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • az

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ai.azure.com

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deploying Scalable Agents loads about 1.6k tokens when it runs. Until then it costs about 255 tokens; SKILL.md has 660 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~255
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/ai-agents-for-beginners at commit 25b7985, republished under its MIT licence (© microsoft). 660 words, ~1,643 tokens.

Download SKILL.mdSave it as .claude/skills/deploying-scalable-agents/SKILL.md (or your agent's skills folder).
name
deploying-scalable-agents
description
Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry. E cover deployment patterns (client-hosted, hosted agents, agent workflows), the agent lifecycle, model routing, response caching, evaluation gates, human-in-the-loop approval, observability with OpenTelemetry, cost optimisation, and smoke-testing deployed agents with the AI Smoke Test action. Based on Lesson 16 of AI Agents for Beginners. USE FOR: deploy one agent go production, scale one agent, Microsoft Foundry hosted agent, Foundry Agent Service, model routing, response caching, evaluation gate, release gate, human approval workflow, agent observability, agent tracing, agent cost optimisation, smoke test one hosted agent, production customer support agent. DO NOT USE FOR: building your first agent (start with Lesson 01), running agents locally on-device (use local-ai-agents / Lesson 17), Azure infrastructure provisioning we no relate to agents, non-Foundry deployment targets.
license
MIT

Deploying Scalable Agents with Microsoft Foundry

Companion skill for Lesson 16 – Deploying Scalable Agents. Use am to help learner move agent from prototype go scalable, observable production deployment. Ground every recommendation inside lesson content and the runnable notebook; no make you invent Foundry APIs.

Triggers

Activate dis skill when learner wan:

  • Deploy agent to Microsoft Foundry as hosted agent and make am versioned/observable.
  • Choose between client-hosted, hosted-agent, and agent-workflow deployment patterns.
  • Add model routing, response caching, or bounded concurrency to control latency and cost.
  • Add evaluation gate so bad agent version no fit ship.
  • Add human-in-the-loop approval step for high-risk actions.
  • Instrument agent wit OpenTelemetry tracing for production observability.
  • Smoke-test deployed agent as quick post-deploy gate.

Core mental model

Production agent na mostly operational skeleton around di model (~80%), no be di model itself. Map every recommendation to one of dis concerns:

ConcernPrototype → Production
Hostingnotebook → versioned hosted service
Identityyour az login → managed identity + scoped RBAC
Statein-memory → externalised thread/memory store
Failuretraceback → retries, fallbacks, alerts
Cost"small small cents" → tracked, routed, cached, budgeted
Qualityeyeballing → automated evaluation gate
Trustyou approve → policy + human-in-the-loop

Deployment patterns (pick one, or combine)

  1. Client-hosted — di reasoning loop dey run inside your process. Max control; you dey own scaling/state.
  2. Hosted agent (Foundry Agent Service) — Foundry dey host the loop, dey store threads, dey enforce RBAC/content safety, dey show di agent for portal. Less control, far less operational wahala.
  3. Agent workflow — multiple agents/tools join body as graph wit branching, approval nodes, and durable checkpoints.

Lifecycle (di loop wey dey deliver agent)

create → version → evaluate (gate) → deploy hosted → observe online → collect failures → repeat. Offline evaluation na gate, no be afterthought — version no go ship unless e clear di threshold. Online observability dey feed real failures back enter di offline test set.

Scaling and cost levers (priority order)

  1. Right-size di model — use di smallest model wey fit pass di evaluation gate.
  2. Route by complexity — small/fast model for simple requests, big model for real reasoning (DIY classifier or Foundry Model Router).
  3. Cache — serve near-duplicate requests without model call.
  4. Stateless design + bounded concurrency — externalise state; retry wit backoff.

Key patterns wey dem suppose reproduce

Point learner to these from di notebook 16-python-agent-framework.ipynb:

  • Request handler: cache → route by complexity → trace span → run → cache.
  • Evaluation gate: score offline test set; return pass_rate >= threshold and only deploy if true.
  • Human approval: @tool(approval_mode="always_require") for actions like large refunds.
  • Tracing: wrap each request inside tracer.start_as_current_span(...) and set attributes like routed.model, customer.id.
Show full SKILL.md (252 more words)Show less

Smoke-testing deployed agent

After deploy, make sure say endpoint really dey answer (green deploy fit still dey silent). Use AI Smoke Test action via .github/workflows/smoke-test.yml with catalog for tests/. Runner dey POST each prompt to POST {project_endpoint}/agents/{agent_name}/endpoint/protocols/openai/responses and e dey check the reply text. Identity need Azure AI User role at Foundry project scope; token audience must be https://ai.azure.com/.

Combine di gates: smoke test (make sure e dey respond, every deploy) → offline evaluation (good enough to ship before promotion) → online evaluation (how e dey perform for real life, continuous).

Enterprise controls

  • RBAC: give every hosted agent managed identity with least privilege.
  • MCP for production: treat every MCP server as untrusted boundary — pin version, scope identity, validate outputs, rate-limit, no ever expose secrets.

Guardrails for the assistant

  • Prefer di canonical FoundryChatClient(...) + provider.as_agent(...) pattern wey dem dey use for di whole course.
  • No promise live-Azure results wey you never comfirm; recommend di smoke-test workflow to confirm deployment.
  • Keep evaluation and cost advice together: evaluation set di quality floor, routing/caching keep cost near dat floor.

<!-- CO-OP TRANSLATOR DISCLAIMER START -->

Disclaimer: Dis document don translate wit AI translation service Co-op Translator. Even tho we dey try make am correct, abeg make you know say automated translation fit get errors or mistakes. Di original document for dia own language na im be di correct source. For important info, make person wey sabi human translation do am. We no go responsible for any misunderstanding or wrong understanding wey fit happen because of dis translation.

<!-- CO-OP TRANSLATOR DISCLAIMER END -->

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in translations/pcm/.agents/skills/deploying-scalable-agents of microsoft/ai-agents-for-beginners.

Open the folder on GitHubat commit 25b7985

Compare with similar skills

Deploying Scalable Agents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deploying Scalable Agents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deploying Scalable Agents this skillmicrosoft/ai-agents-for-beginners77k—~1.6kAutomated safety check: PassMIT
Aspiremicrosoft/aspire.dev1964 repos~1.1kAutomated safety check: PassMIT
Aspire MonitoringCommunityToolkit/Aspire629—~3.5kAutomated safety check: PassMIT
Temps Best Practicesgotempsh/temps826—~2.9kAutomated safety check: PassApache-2.0
Backdoor Deploymentmicrosoft/Docker-Provider174—~7kAutomated safety check: PassCustom licence
Observability Architecturemajiayu000/litellm-rs117—~1.3kAutomated safety check: PassMIT

Similar skills

  • Aspire

    microsoft/aspire.dev

    Official

    Orchestrates Aspire distributed applications using the Aspire CLI for running, debugging, and managing distributed apps.

    196 GitHub starsUsed in 4 repos~1.1k tokens
    DevOps & CloudAuto-check passed
  • Aspire Monitoring

    CommunityToolkit/Aspire

    ANALYSIS SKILL - Observe Aspire apps: logs, traces, metrics, resource state, telemetry export, browser telemetry, and the standalone dashboard.

    629 GitHub stars~3.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Temps Best Practices

    gotempsh/temps

    Best-practices reference for preparing and instrumenting applications on Temps.

    826 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Backdoor Deployment

    microsoft/Docker-Provider

    Official

    Validate a container image change via backdoor deployment. An agent skill from microsoft/Docker-Provider.

    174 GitHub stars~7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Observability Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

    117 GitHub stars~1.3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Azure Monitor OpenTelemetry Exporter for Java. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 6 repos~2.2k tokens
    DevOps & CloudAuto-check passed

More from microsoft/ai-agents-for-beginners

All 123 skills in this repo
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    A skill your agent uses when the user asks to create, scaffold, or edit Jupyter notebooks (.ipynb) for experiments, explorations, or tutorials; prefer the bundled templates and run the helper script…

    77k GitHub starsUsed in 8 repos~1k tokens
    Auto-check passed
  • Azure Openai To Responses

    microsoft/ai-agents-for-beginners

    Official

    Migrate Python apps from Azure OpenAI Chat Completions to the Responses API.

    77k GitHub stars~6k tokensUpdated 19 days ago
    Auto-check: notes
  • Azure Openai To Responses

    microsoft/ai-agents-for-beginners

    Official

    Shift Python apps dem from Azure OpenAI Chat Completions go Responses API.

    77k GitHub stars~6k tokensUpdated 19 days ago
    Auto-check: notes
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    Kasuta, kui kasutaja palub luua, üles ehitada või redigeerida Jupyteri märkmikke (.ipynb) katsetuste, uurimiste või juhendite jaoks; eelista kaasasolevaid malle ja käivita abiskript newnotebook.py…

    77k GitHub stars~1.2k tokensUpdated 19 days ago
    Auto-check passed
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    Käytetään, kun käyttäjä pyytää luomaan, alustamaan tai muokkaamaan Jupyter-muistikirjoja (.ipynb) kokeita, tutkimuksia tai opetusohjelmia varten; käytä mieluummin mukana olevia mallipohjia ja…

    77k GitHub stars~1.3k tokensUpdated 19 days ago
    Auto-check passed
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    À utiliser lorsque l'utilisateur demande de créer, structurer ou modifier des notebooks Jupyter (.ipynb) pour des expériences, explorations ou tutoriels ; privilégiez les modèles fournis et exécutez…

    77k GitHub stars~1.4k tokensUpdated 19 days ago
    Auto-check passed

Questions about Deploying Scalable Agents

What does Deploying Scalable Agents do?

Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry. Deploying Scalable Agents is an agent skill from microsoft/ai-agents-for-beginners, published by the product's own GitHub organization. Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry.

When should I use Deploying Scalable Agents?

Deploying Scalable Agents fits situations like: : deploy one agent go production; scale one agent; microsoft Foundry hosted agent; foundry Agent Service.

How do I install Deploying Scalable Agents in Claude Code?

Run `npx skills add microsoft/ai-agents-for-beginners --skill deploying-scalable-agents -a claude-code`. Or copy the skill folder (translations/pcm/.agents/skills/deploying-scalable-agents in microsoft/ai-agents-for-beginners) into .claude/skills/deploying-scalable-agents in your project. Claude Code loads it when a task matches its description.

How do I install Deploying Scalable Agents in Codex?

Run `npx skills add microsoft/ai-agents-for-beginners --skill deploying-scalable-agents -a codex`. Or copy the skill folder (translations/pcm/.agents/skills/deploying-scalable-agents in microsoft/ai-agents-for-beginners) into .agents/skills/deploying-scalable-agents in your project. Codex loads it when a task matches its description.

Can I use Deploying Scalable Agents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/ai-agents-for-beginners --skill deploying-scalable-agents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deploying-scalable-agents, .gemini/skills/deploying-scalable-agents, .github/skills/deploying-scalable-agents and .opencode/skills/deploying-scalable-agents in your project.

What does Deploying Scalable Agents need to run?

Going by SKILL.md and its folder, Deploying Scalable Agents needs the command-line tools its instructions call (az).

Does Deploying Scalable Agents access the network?

SKILL.md names 2 domains. In commands or code: ai.azure.com; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Deploying Scalable Agents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deploying Scalable Agents use?

Deploying Scalable Agents is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deploying Scalable Agents use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deploying Scalable Agents?

Skills that share tags, products or a category with Deploying Scalable Agents: Aspire (microsoft/aspire.dev, 196 stars), Aspire Monitoring (CommunityToolkit/Aspire, 629 stars), Temps Best Practices (gotempsh/temps, 826 stars) and Backdoor Deployment (microsoft/Docker-Provider, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deploying Scalable Agents?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/ai-agents-for-beginners, which has 76,612 GitHub stars. The repository holds 123 skills in this directory. The repository was last updated on September 19, 2026.

Source: microsoft/ai-agents-for-beginners on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.