Agent skill

Ak Dev Testing Conventions

by yaalalabs in yaalalabs/agent-kernel

Testing conventions, patterns, and automation for Agent Kernel development.

Apache-2.0Auto-check passedTesting & QA

Install Ak Dev Testing Conventions

skills CLI
$ npx skills add yaalalabs/agent-kernel --skill ak-dev-testing-conventions -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yaalalabs/agent-kernel ak-dev-testing-conventions --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yaalalabs/agent-kernel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ak-dev-testing-conventions .claude/skills/ak-dev-testing-conventions && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ak-dev-testing-conventions
GitHub stars
191
Token cost
~14k tokens
SKILL.md length
5,263 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Testing conventions, patterns, and automation for Agent Kernel development.

  • Works in 6 steps: Use ordered tests… → Use session-scoped fixtures for test… → Mock external services (LLM APIs, cloud… → …
  • Writing tests for new features
  • SKILL.md covers Running Tests, Test File Organization, Test Patterns and Built-in Test Framework, plus 3 more sections
  • Calls uv, terraform and helm; needs GITHUB_TOKEN and OPENAI_API_KEY

What it does

Ak Dev Testing Conventions is an agent skill from yaalalabs/agent-kernel. Testing conventions, patterns, and automation for Agent Kernel development. Use this skill when writing tests for new features, debugging test failures, or understanding the test infrastructure. Covers pytest patterns, async testing, mocking, the built-in Test framework, and CI/CD test workflows.

Its SKILL.md is about 14k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Unit testing, Failing and flaky tests and CI/CD. It works with pytest. The repository describes itself as: The Operating System for Scalable Enterprise AI Agents - Run, orchestrate, and deploy Compliant Enterprise AI Agents at scale across frameworks, without lock-in, rewrites or… The licence is Apache-2.0.

When your agent uses it

  • Writing tests for new features
  • Debugging test failures
  • Understanding the test infrastructure

Example prompts

  • “/ak-dev-testing-conventions”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Use ordered tests (@pytest.mark.order(n)) when testing conversational flows where follow-up questions depend on prior context
  2. Use session-scoped fixtures for test clients that are expensive to create
  3. Mock external services (LLM APIs, cloud services) in unit tests — only hit real APIs in integration tests
  4. Test both success and failure paths — especially for hooks and guardrails
  5. Use DummyAgent/DummyRunner to isolate the component under test from framework-specific behavior
  6. Test session persistence — verify that state survives across multiple Runtime.run() calls within the same session

What it can do on your machine

Read from SKILL.md and the folder at commit e03a602. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • terraform
    • helm
    • aws
    • az

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, helm, aws and az, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GITHUB_TOKEN
    • OPENAI_API_KEY
    • APP_PRIVATE_KEY
    • COPILOT_REQUEST_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ak Dev Testing Conventions loads about 14k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 5,263 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from yaalalabs/agent-kernel at commit e03a602, republished under its Apache-2.0 licence (© yaalalabs). 5,263 words, ~13,745 tokens.

Download SKILL.mdSave it as .claude/skills/ak-dev-testing-conventions/SKILL.md (or your agent's skills folder).
name
ak-dev-testing-conventions
description
Testing conventions, patterns, and automation for Agent Kernel development. Use this skill when writing tests for new features, debugging test failures, or understanding the test infrastructure. Covers pytest patterns, async testing, mocking, the built-in Test framework, and CI/CD test workflows.
license
Apache-2.0
metadata.author
yaalalabs
metadata.category
developer

Testing Conventions

Running Tests

bash
cd ak-py
uv run pytest                           # Run all tests with coverage
uv run pytest tests/test_runtime.py     # Run specific test file
uv run pytest -k "test_session"         # Run tests matching pattern
uv run pytest -x                        # Stop on first failure

Coverage and HTML reports are auto-generated per pyproject.toml:

toml
[tool.pytest.ini_options]
addopts = "--cov=src --cov-report=term --cov-report=html --html=report.html"

Test File Organization

Tests live in ak-py/tests/ and follow the naming convention test_<module>.py.

Three subsystems ship a reusable contract suite a new backend must be run against, rather than per-backend assertions written from scratch: QueueTransportContract (pipeline/testing.py), SandboxProviderContract (sandbox/testing.py), and KnowledgeBaseContract / DocumentStoreContract (knowledgebase/testing.py). All three ship next to the ABC they constrain, so a bring-your-own backend author can import and subclass them out of tree; each module imports pytest and is therefore kept out of its package's exports. In all three cases the contract classes are deliberately not named Test*, so pytest collects them only through the subclasses that supply a fixture.

Test FileTests
test_base.pySession, Agent, Runner abstractions; Agent.run_options / RESERVED_RUN_OPTIONS / validate_run_options and Runner._native_kwargs (#754); the RunOptionsFactory export, Agent.run_options_factory and resolve_run_options (no-factory copy, sync and async merge, argument identity, non-mapping TypeError, reserved-key ValueError, propagation, static dict untouched) (#758)
test_runtime.pyRuntime registration, execution, hooks
test_stream_events.pycore/event.py's StreamEvent discriminated union: every member round-trips through JSON (type discriminator), rejects an unknown type, and stays JSON/pickle-safe (no framework-native fields)
test_runtime_stream_events.pyRuntime.stream()'s streaming contract (specs docs/specs/523-ag-ui-support/ and docs/specs/670-streaming-post-hooks/): a bare str from an unmigrated runner fails loudly as a pydantic ValidationError, every event reaches PostHook.on_stream_event(), a returned list emits N chunks and ends the chain, a single return of a different type raises TypeError, a hook returning None drops the whole chunk, delta is populated only for TextDelta and taken from the emitted event, StreamHalt closes open boundaries then yields one error chunk and stores no session, any other exception propagates, and the final chunk is a bare StreamChunk(done=True)
test_stream_boundaries.pyStreamBoundaryTracker (core/stream.py) directly: open/close pairing per kind, innermost-first drain order, drain() clearing, and the two malformed-sequence cases it tolerates (closing an id never opened, opening one twice)
test_module.pyModule load/unload, wrapping; Module.run_options merge, chaining, reserved-key and unloaded-agent errors, _native_agent_name override, the base-class pre_hook / post_hook resolving through _native_agent_name, and the guard that a subclass implementing only _wrap and load constructs (#754); the positional factory parameter (stored, replaced, beside keywords, agent= still accepted, a non-callable or a second positional factory raising TypeError and storing nothing, a keyword named factory treated as an option) (#758)
test_session.pySession state, caches, context vars
test_session_cache.pyLRU SessionCache
test_sessions_in_memory.pyInMemorySessionStore
test_sessions_redis.pyRedisSessionStore missing-config error, shared RedisDriver retry exhaustion
test_sessions_valkey.pyValkeySessionStore round trips (fake client), shared ValkeyDriver retry exhaustion
test_sessions_dynamodb.pyDynamoDBSessionStore Binary wrap/unwrap, missing-item skip (mocked driver)
test_shared_drivers.pyShared DB drivers (core/util/driver/): retry scope, ping/reconnect, command surface, DynamoDB item-dict semantics, S3Driver object operations and its no-probe-on-connect contract
test_multimodal_redis_store.pyRedisAttachmentStore index TTL refresh, JSON round trip, pruning (mocked driver)
test_multimodal_source_forms.pyMultimodalPreHook attachment source-form classification (spec #523 §8): bare base64 and base64 data: URIs are described/stored/stripped; http(s):///s3:// and non-base64 data: URIs are retained undescribed; empty data: payloads are dropped
test_thread_source_forms.pyThe same source forms through ConversationThreadManager.store_attachments (issue #669): base64 stored as bytes and replaced by an AgentRequestAttachmentRef; a remote reference recorded by url with empty data and its request passed through unreplaced; mixed/multiple attachments keeping order; plus an end-to-end guard that a URL image survives MultimodalPreHook to the agent
test_multimodal_tools.pyanalyze_attachments content routing per stored-record shape: a remote record's url is sent as-is (an image as image_url, anything else named in text) and never wrapped as base64; a stored record's base64 becomes a data: URI, a PDF a file part, any other type a [Document: ...] line; plus the two early returns that never reach the store or the LLM
test_config.pyAKConfig loading, env vars
test_test_config.pyAKTestConfig (Test framework config) loading, defaults
test_tool.pyToolContext, cache
test_tool_openai.pyOpenAI ToolBuilder
test_tool_crewai.pyCrewAI ToolBuilder
test_tool_langgraph.pyLangGraph ToolBuilder
test_tool_adk.pyGoogle ADK ToolBuilder; its runner-driving mock agents come from the file's _mock_agent() helper (#758)
test_tool_smolagents.pySmolagents ToolBuilder
test_tool_pydanticai.pyPydantic AI ToolBuilder
test_openai_runner.pyOpenAIRunner execution, error handling; run options forwarded to Runner.run / run_streamed, and a real RunHooks resolving Session.current() / Agent.current() through Runtime.run and Runtime.stream (#754); a factory-merged mapping wins at Runner.run / run_streamed, resolution is awaited once per method, a failing resolution becomes the error reply in run and propagates from stream, and a real factory declared through OpenAIModule.run_options runs per call (#758). Every mock agent handed to a runner comes from a _mock_agent() helper carrying run_options and an AsyncMock resolve_run_options returning it, the contract all runner test files share since #758
test_adk_runner.pyGoogleADKRunner execution; reserved keys, RUNNER_CONSTRUCTOR_OPTIONS splitting constructor-bound options from run_async-bound ones, constructor options reaching the per-run Runner(...), run forwarding a caller RunConfig untouched, and stream copying a caller RunConfig to force streaming_mode=SSE with a once-per-runner warning (#754); _split_run_options and _setup_session_context taking the resolved mapping, run handing a factory-merged mapping to the setup (a factory plugins reaching the constructor) and forwarding its run_config, stream SSE-copying a factory run_config, resolution awaited once per method, and an empty request returning before resolving (#758)
test_crewai_runner.pyCrewAIRunner execution (mocked Crew kickoff); reserved keys, declared options reaching the per-run Crew(...) with verbose overridable, empty options reproducing today's kwargs, and the module resolving the agent by role and rejecting reserved keys (#754); a factory-merged max_rpm and verbose reaching Crew(...), resolution awaited once, and a rejected request returning before resolving (#758)
test_smolagents_runner.pySmolagentsRunner execution, multimodal requests, error handling; reserved keys, max_steps forwarded beside reset, a bypassed reset/additional_args overwritten by the AK-owned values, additional_args dropped when there is no framework context, and the module resolving the agent by its native name (#754); a factory-merged max_steps reaching agent.run beside reset, resolution awaited once, and a rejected request returning before resolving (#758)
test_pydanticai_runner.pyPydanticAIRunner execution, structured output, BinarySerde session round-trip, multimodal wiring; reserved keys, declared options forwarded to run beside the AK-owned keys, stream dropping event_stream_handler with one warning while keeping the declared dict intact, and reserved keys rejected at declaration through the module (#754); a factory-merged usage_limits reaching run, the stream filter applied to the resolved mapping (_stream_run_options taking a mapping and leaving it untouched), resolution awaited once per method, and an empty request returning before resolving (#758)
test_langgraph_runner.pyLangGraphRunner execution; _merge_run_config deep-merge semantics (thread_id wins inside configurable, list-valued keys like callbacks concatenate AK-first, a scalar AK value wins over a None caller config, neither input is mutated), run/stream forwarding the merged config and other options to ainvoke/astream_events, the declared options surviving unmutated across runs, reserved keys, the nested config.configurable.thread_id and top-level RunnableConfig keys rejected at declaration, and aget_state receiving the merged config (#754); a factory-merged config deep-merged in run and stream with aget_state reading the same merged config back, resolution awaited once per method, and a rejected request returning before resolving (#758)
test_trace_langfuse_langgraph.pyLangFuse LangGraph trace runner; a caller-supplied callbacks list is appended after the Langfuse handler rather than replacing it (#754); a factory-merged config reaching ainvoke through the traced runner with the Langfuse handler kept first (#758)
test_langgraph_reasoning_live.pyLangGraph reasoning against a REAL reasoning model, env-gated (AK_TEST_REASONING_MODEL; skipped in normal runs). Guards the premise the chunk-feeding unit tests cannot: that the model streams a summary at all (it must be asked — reasoning={"summary": "auto"}) and that LangChain surfaces it under content_blocks. Builds a bare StateGraph rather than create_react_agent, since the test needs only one model call and a ReAct loop would add nothing
test_guardrail.pyGuardrail factories, hooks
test_api_http.pyREST API handler
test_api_mcp.pyMCP.get_http_app(): mcp.stateless_http is passed to http_app() (fastmcp 3 removed the FastMCP(stateless_http=...) constructor kwarg, so the constructor must stay name-only), and a second call reuses the built server while still applying the configured mode
test_chat_service_core.pyChatService execution core (execute/execute_stream): typed replies, prebuilt request lists, validation, error propagation, wrapper wire shapes
test_chat_service_streaming.pyChatService SSE/stream chunk formatting
test_integration_adapter_contract.pyIntegrationAdapterContract subclassed once per built-in messaging adapter: identifier resolution, ignorable deliveries, reply-context budget, queue round trip (start here for a new platform)
test_slack_integration.pySlack adapter: event -> InboundRequest and reply -> Slack messages, plus Bolt's signed dispatch through the webhook host (pattern for per-platform adapter tests)
test_integration_roundtrip.pyA platform event through the whole in_memory topology to a recording outbound adapter
test_whatsapp_integration.pyWhatsApp adapter: webhook delivery -> InboundRequest, and agent reply -> Cloud API sends
test_teams_integration.pyTeams adapter: Bot Framework activity -> InboundRequest (through process_activity), and agent reply -> proactive continue_conversation, including the per-delivery BotFrameworkAdapter
test_gmail_integration.pyGmail adapter: unread mail -> InboundRequest, and agent reply -> a threaded reply
test_messenger_integration.py, test_instagram_integration.pyMessenger / Instagram adapters: webhook delivery -> InboundRequest, and agent reply -> Send API calls
test_telegram_integration.pyTelegram adapter: update -> InboundRequest, and agent reply -> Bot API sends
test_integration_adapter_factory.pyIntegrationAdapterFactory: how an integration attribute value becomes an outbound adapter — bare names, per-platform dotted-path overrides, instance caching, and the configuration errors (unknown name, wrong type, unimportable path, missing extra); an inbound adapter is never resolved by name
test_integration_producer.pyIntegrationProducer: what a parsed platform message looks like on the input queue — routing attribute, reply_-prefixed context and its 8 KB budget, ordering/dedup keys from the platform ids, prebuilt request list in the body
test_integration_webhook_handler.pyWebhookRESTRequestHandler: the generic host for a push-based inbound adapter — route mounting, batched deliveries, edge acknowledgement, ignored deliveries, SDK-owned responses, verification rejection, and RESTAPI.run refusing it
test_integration_poller_runner.pyPollerRunner: hosting a pull-based inbound adapter — enqueue-then-mark_handled ordering, a failing poll costing one iteration, shutdown within one interval, and the topology rejections both ways
test_attachment_offload.pyAttachmentStorageManager.offload: image/file requests rewritten to refs in place, other requests kept in order, the caller's message when multimodal is disabled, session_cache refused for an attachment but not for a text-only list
test_thread_integration.pyThread integration: ThreadRecorder ordering/enforcement, AgentThreadRequestHandler recording + no-phantom-thread prechecks, stream accumulation, end-to-end read-back
test_thread_pipeline_recording.pyThread recording split across the queue: ThreadRequestHandler marking the message and committing the user message (and offloading attachments) before enqueue, its rejections leaving no phantom thread (missing user_id, unavailable agent, ASYNC mode, STREAM without a chunk-streaming store), deferred requests staying unmarked, AgentRunner/StreamAgentRunner appending the reply only for a marked message and only after the output send, a thread-store failure never retrying the run, IOHandler.run(request_handler=...) replacing rather than joining the chat route and refusing a non-RequestHandler, and the end-to-end read-back
test_thread_router.pyThread read routes (ThreadRESTRequestHandler): pagination, Authoriser 401/403 semantics
test_authoriser_shared.pyShared Authoriser in agentkernel.auth: package-export identity, guard that the thread package no longer exposes it, AuthValidatorAuthoriser adapter, AuthorisedRESTRequestHandler inheritance
test_akagentrunner_stream.pyServerless ServerlessStreamAgentRunner (SQS streaming)
test_serverless_agent_runner_schedule.pyServerless runners' trigger consumption: request_id/user_id body fallback, attribute precedence, missing-in-both error path
test_akresponsehandler.pyServerless response handler (CHAT_RESPONSE / STREAM_CHUNK broadcast)
test_ws_lambda_stream.pyWebSocket Lambda router in stream mode
test_cli_tester.pyCLI test framework
test_auth_handler.pyAuth handler
test_akauthorizer.pyAWS Lambda authorizer
test_lambda_router.pyLambda routing
test_sqs_handler.pyAWS SQSHandler config, client, message sending
test_serverless_request_handle.pyBaseRequest/BaseRunRequest parsing from serverless payloads
test_firestore_database_id.pyShared FirestoreDriver (core/util/driver/firestore.py, explicit constructor params) named database_id configuration
test_ak_logger.pyAKLogger level resolution, configuration
test_error_util.pyuser_facing_error_message error mapping
test_thread_runner.pyThreadRunner task validation, failure/shutdown semantics
test_ecs_sqs_consumer_parallel.pyECSSQSConsumer message processing + delete/retry semantics
test_ecs_agent_runner_schedule.pyECS runner trigger consumption: request_id/user_id body fallback with attribute precedence, and ChatService's status forwarded to the output queue instead of discarded
test_ecs_output_consumer_status.pyECS output consumer persisting status_code on stored records (default 200, permanent failure 500)
test_deployment_queue_contracts.py#495 public-interface cleanup: pipeline.transport (QueueTransport/QueueMessage) is the only public queue API; RawQueueConsumer + SQSHandler's send models are internal; removed public names (QueueHandler, QueueConsumer, deployment.common.queue_*) raise ImportError
test_pipeline_agent_runner.pyAgentRunner/StreamAgentRunner: chat execution via ChatService, reply forwarding with STATUS_CODE attribute, per-chunk dedup suffixes, run() rejecting in_memory transport
test_pipeline_agent_runner_schedule.pyPipeline runners' trigger consumption: request_id/user_id resolved from the message body, attribute precedence, and body-resolved metadata injected back into the attributes for output forwarding
test_pipeline_bookkeeping.pyDelivery bookkeeping for transports lacking native receive counts/dedup (spec #495 §6): InMemoryBookkeepingStore/RedisLikeBookkeepingStore attempt counters, retry-safe dedup claims, BookkeepingStoreFactory backend selection
test_pipeline_sqs_transport.pySQSTransport: send/fetch/ack/nack/dead_letter, graceful shutdown handling, fetch-wait slicing, shared wire-format primitives
test_pipeline_kafka_transport.pyKafkaTransport against a fake in-memory Kafka cluster: send/fetch/ack/nack, dead-letter routing scoped by topic, delivery errors, consumer capacity check
test_pipeline_nats_transport.pyNatsTransport against a fake JetStream behind a real _NatsLoop bridge: subject/header construction, stable client-side (crc32) partition hashing, num_delivered mapping, nak redelivery, term() on permanent failure, one-in-flight-per-partition, stream-scoped dedup, auto_provision create-vs-verify, consumer capacity warning, and the full QueueTransportContract with no skips
test_response_store_in_memory.pyInMemoryResponseStore: get_record (status_code exposed), add_chunk/stream chunk-streaming for local SSE
test_transport_contract.pyThe reusable QueueTransportContract (pipeline/testing.py) run against the in_memory transport
test_transport_contract_live.pyThe same contract against REAL brokers, env-gated (AK_TEST_NATS_URL / AK_TEST_KAFKA_BOOTSTRAP; skipped in normal runs) with per-test unique streams/topics; run in CI by test-reusable.yaml's transport-integration-tests job over the transport examples' compose stacks. Documents two live-only timing traps: the per-partition pull window must stay below ack_wait, and partition counts are chosen from the real partitioner mappings (crc32 / murmur2)
test_pipeline_request_handler.pyPipeline RequestHandler over FastAPI TestClient: rest_sync parity (stored status_code honored), rest_async accept/poll, SSE streaming end-to-end, multipart-on-in_memory only, and SSE error frames carrying handler-owned text only — a store/driver TimeoutError/Exception message never reaches the client (CodeQL py/stack-trace-exposure)
test_pipeline_response_handler.pyPipeline ResponseHandler delivery paths: REST records, the USER_ID-presence WS routing marker, STREAM_CHUNK/CHAT_RESPONSE pushes, missing-attribute retries, permanent-failure frames/records
test_pipeline_io_handler.pyIOHandler.run() topology validation and fail-fasts (ASYNC-on-in_memory without a validator, broker WS modes without a push token, broker + non-shared response store), signal handlers, graceful-drain exit code
test_pipeline_ws.pyThe gateway tier: LocalConnectionRegistry, PodPushWebSocketHandler store-lookup delivery + stale-connection cleanup, /internal/push auth, the native /ws route (1008 closes, chat enqueue attributes, custom routes), WebSocketGateway fail-fasts, single-process ASYNC/STREAM end-to-end over in_memory, and cross-"pod" delivery between two gateway apps sharing one connection store
test_session_connection_store.pyThe WSConnectionStore contract over the in-memory/redis-like/DynamoDB implementations, plus SessionStore.get_connection_store() per backend (store-less backends raise actionably)
test_sandbox.pySandbox core: model/capabilities, error hierarchy, config, provider contract, manager + factory + embedded broker, agent surface (system tools + task-completion pre-hook), agents scoping
test_sandbox_broker.pyBroker flavors (embedded/thread) end-to-end, thread loop-identity contract, wait-policy promotion + late-completion recovery (including the bounded result_summary outcome on check_sandbox_task), completion ingestion
test_sandbox_providers.pylocal_subprocess (real subprocess), docker (mocked SDK), e2b/daytona (mocked cloud SDKs), ec2_ssm (mocked boto3), and kubernetes (fake SDK with a functional tar/base64 exec protocol; manifest shapes, NetworkPolicy gating, RBAC-impersonation headers per call shape) providers, run against the reusable SandboxProviderContract
test_sandbox_queue_broker.py#503 queue broker: BrokerWireCodec binary round trips, error_type stamping, removed config fields, QueueExecutionBroker client (bounded waits, promotion, typed re-raise, ceiling/size guards, FIFO destroy), QueueBrokerWorker two-loop model (output-queue delivery, truncation, split permanent-failure hooks, sweep inventory), and manager-level wait-then-check_sandbox_task recovery over in_memory
test_pipeline_factory_seams.py#503 explicit-config seams: QueueTransportFactory/ResponseStoreFactory with an explicit block never read AKConfig (asserted loudly), default paths unchanged, response-store ttl override
test_response_store_scan.pyThe optional ResponseStore key-scan capability (supports_key_scan/scan_records) across the four built-ins, ABC default opt-out
test_authoriser_shared.pyShared Authoriser in agentkernel.auth: AuthValidatorAuthoriser adaptation, export identity, and the guard that the thread package no longer re-exports it
test_schedule_model.pyScheduleSpec one-of/timezone/session_mode validation, chat-envelope parsing, ScheduledTask JSON round trip (JSON primitives only)
test_schedule_manager.pyScheduleManager: get() gating + singleton, provider/transport fail-fast, semantic validation matrix (including the named-agent precheck and the unnamed-agent exemption), create ordering + rollback, trigger-body freezing, ownership, amendment rules (occurrence rule replaced as a unit, untouched when the amendment names none of it), cancellation, occurrence recording (never raises)
test_schedule_store.pyEvery ScheduleStore backend against the one contract the in_memory class pins — in_memory, redis/valkey through a fake redis-like client injected as store._driver._client, dynamodb through a fake driver whose table.scan replays LastEvaluatedKey pages — plus index cleanup on delete, TTL behaviour (none by default), and ScheduleStoreBuilder built-in/BYO/unknown-type/missing-extra resolution
test_schedule_provider_local.pyLocalScheduleProvider: next-fire computation, one-time vs re-armed occurrences, token substitution, body-only delivery into InMemoryTransport, pause/delete disarm, ScheduleProviderFactory resolution
test_schedule_provider_eventbridge.pyEventBridgeScheduleProvider against a mocked boto3.client: cron/at expression translation, exact registration/amendment/removal kwargs sent to the Scheduler API, error mapping to ScheduleError, ScheduleProviderFactory wiring
test_schedule_router.pyScheduleRESTRequestHandler: 404 when unconfigured, the three 401 variants from the shared AuthorisedRESTRequestHandler, listings forced to the authorised user, 403-before-404 ordering, PUT amendment happy path + validation 400s, DELETE returning the cancelled task, mount-time validation of the configured backends (get_router() building the manager, an unusable provider failing the mount, mounting while unconfigured still allowed), and the guard that importing the package does not pull in FastAPI
test_secret_manager.pySecretManager: env → cache → provider precedence (empty env var is a miss, env hits never cached), SecretNotFoundError vs default, provider SecretError never masked, os.environ never written, key validation, TTL expiry with a fake clock, invalidate/clear, no lock held across the provider call, current() singleton built from config
test_secret_factory.pySecretProviderFactory: env/aws_ssm/dotted-path resolution, unknown names and non-subclass paths rejected, missing aws extra raised before the prefix check, never reads AKConfig
test_secret_providers.pyEnvSecretProvider and AWSSMSecretProvider (mocked boto3: path composition, ParameterNotFound → None, other errors → SecretError, prefix validation), each run against the reusable SecretProviderContract (secret/testing.py)
test_schedule_tools.pySchedule system tools: SystemToolFactory registration + schedule.agents scoping, prompt-suffix content, disabled short-circuit, acting-user read from the session volatile cache, and each tool's JSON contract including the no-identity and unknown-agent errors
test_chat_service_schedule.pyChatService interception on all four entry points: 202 wire shapes, streaming terminal chunk, unconfigured 400, occurrence recording, no scheduling field leaking as AgentRequestAny
test_serverless_status_propagation.pyThe status a chat run produced (e.g. 202 deferred-to-schedule, 4xx rejected) survives the serverless queue round trip: ServerlessAgentRunner forwards it as the ATTR_STATUS_CODE output-message attribute, ResponseHandler stores it on the response record, and the REST surface (DefaultEndpointsHandler, response-store polling) replays it to the caller
test_pipeline_agent_runner_schedule.pyPipeline runners' trigger contract: request_id/user_id body fallback, attribute precedence, attribute injection for output forwarding
test_ecs_agent_runner_schedule.pyECS runners' body fallback (_get_record_attributes with/without a parsed body) and status_code custom attribute
test_ecs_output_consumer_status.pyECSOutputConsumer stores status_code (present / absent → 200 / permanent failure → 500)
test_serverless_agent_runner_schedule.pyServerless runners' body fallback in both _get_record_attributes implementations
test_knowledgebase_model.pyKnowledgeCapabilities defaults and model_dump() shape, the record TypedDicts, and KnowledgeBase.validate_capabilities's two invariants (unreachable declaration, query ⇔ query_language) — asserting a bare capabilities object never validates on its own
test_knowledgebase_base.pyThe reshaped ABC: three abstract members, the five capability-gated operations raising KnowledgeCapabilityError, the guard that the ABC exposes no read, schema() merge order with capabilities written last and unoverridable, _derived_schema(), and the fetch-gated [<id>] prefix in format_results()
test_knowledgebase_builder.pyKnowledgeBuilder gating: the tool matrix and emission order, search_kb's any-backend-declares-search gate, write_kb's always-emitted/per-call check and build(writable=False) withholding it, read_kb's routing between query() and search() (moved here off the ABC), capability mismatches returning an actionable string rather than raising, semantic_map resolution across queries/browse paths/fetch id segments, and the generic query/params write metadata
test_knowledgebase_okf_roles.pyThe OKF role model: OKFRoleRegistry.from_config over literal dicts (single role, multi-database, a different role per database), OKFRole.writable as the one place the consumer-vs-writable distinction lives, may_write/writable_databases_for/has_any_role, and both AKConfigErrors — a database naming no agent, and one agent as both producer and curator of the same database — each naming the database
test_knowledgebase_okf_config.pyThe okf block and OKFCapabilityManager: YAML and AK_OKF__* parsing, AKConfig.okf defaulting to None and a pre-change config file parsing unchanged, every field default, _resolve_store's three branches and both type/uri agreement failures, the read-only-store warning, and the laziness guarantees (validation walks no store; a backend walks one on first use)
test_knowledgebase_okf_prompts.pyOKFPromptComposer.compose from config alone with no backend constructed: the per-database line and its access marker, both mandates for an agent holding different roles in two databases, each mandate scoped under its own database, '' for an unnamed agent, and the navigation protocol present exactly once regardless of database count
test_knowledgebase_okf_tools.pyThe OKF agent surface end to end over real bundles in tmp_path: the tool set per role, func.__name__ == SystemTool.name, read scoping falling out of the per-agent builder, the per-(agent, database) write refusal as a string, no tool carrying the prompt section (it arrives through SystemToolFactory.get_prompt_sections), SystemToolFactory wiring, and write permission keyed to the tool's owner across a handoff (a stale entry-agent ToolContext is ignored)
test_knowledgebase_backends.pyThe three SDK-backed backends with their clients monkeypatched: the read→search/query renames, each declaration, Neo4j's generic-with-cypher_*-fallback write metadata, and Starburst's db_schema no longer shadowing schema()
test_knowledgebase_stores.pyDocumentStoreContract over a real tmp_path and a fake boto3 client, plus the containment matrix (.., absolute, normalising escapes, symlinks) and the global-lexicographic list() ordering case (a/z.md before ab/b.md)
test_knowledgebase_okf_parser.pyOKFParserUtil: frontmatter splitting, every diagnostic code reachable, trust derived from verified alone, staleness against an injected now, link extraction, the bounded field_tokens index, and a guard that the module pulls in no store or HTTP client
test_knowledgebase_okf_manager.pyOKFManager end to end: manifest walk and truncation, ranking determinism, browse-at-a-namespace with and without a curated index.md, write-through visibility, one walk under two concurrent boundary-crossing callers, and that nothing is ever filtered on trust or staleness
test_knowledgebase_contract.pyThe reusable KnowledgeBaseContract (knowledgebase/testing.py) run against FakeKnowledgeBase in four capability shapes, OKFManager over a real local bundle, and the three SDK backends with their clients monkeypatched. The schema()-is-callable assertion is the Starburst-collision regression guard
test_knowledgebase_okf_envelope.pyThe declared scale envelope: a generated 10,000-concept bundle keeps all 10,000, and tracemalloc allocations attributable to the manifest walk stay under a 25 KB per-concept budget — measured over a 2,000-concept slice, 250 MB projected at 10,000 (allocations, not RSS)
test_knowledgebase_exports.pyEvery agentkernel.knowledgebase.__all__ name resolves through the PEP 562 lazy map, and chromadb/neo4j/trino/boto3 stay out of sys.modules on import — the gate a new export has to pass
test_factory.pyShared pluggable-backend helpers (resolve_dotted, require_extra, AKConfigError) in core/util/factory.py
test_store_builders.pySession/thread/multimodal store builders: fail-loud on unknown type, BYO dotted-path subclass resolution
test_trace.pyTrace factory built-in resolution, BYO dotted path, unknown-type error
Show full SKILL.md (1,917 more words)Show less

Test Patterns

Dummy Implementations for Unit Testing

Create minimal implementations of abstract classes:

python
from agentkernel.core.base import Agent, Runner, Session
from agentkernel.core.model import AgentReplyText, AgentRequest, AgentRequestText


class DummyRunner(Runner):
    async def run(self, agent, session, requests):
        prompt = requests[0].prompt if isinstance(requests[0], AgentRequestText) else ""
        return AgentReplyText(response=f"ok:{prompt}")

    async def stream(self, agent, session, requests):
        # Runner.stream() is abstract — implement even in test doubles, and yield `StreamEvent`
        # members (never a bare `str` — `StreamChunk.event` rejects it with a `ValidationError`).
        # Raise NotImplementedError() (with a trailing `yield`) if the test doesn't exercise streaming,
        # or yield MessageStart/TextDelta/MessageEnd (own the message's boundaries yourself) to test
        # Runtime.stream() / AgentService.stream_multi().
        raise NotImplementedError()
        yield


class DummyAgent(Agent):
    def __init__(self, name="test-agent"):
        runner = DummyRunner("DummyRunner")
        super().__init__(name, runner)

    def get_description(self) -> str:
        return "Test agent"

    def get_a2a_card(self):
        return None

A double that does exercise streaming yields events, not strings. Runner.stream() is typed AsyncGenerator[StreamEvent, None], and StreamChunk.event is a discriminated union — so a bare str raises a ValidationError the moment Runtime.stream() wraps it. A double owns its own boundaries:

python
from agentkernel.core.event import MessageEnd, MessageStart, TextDelta


class StreamingDummyRunner(Runner):
    async def run(self, agent, session, requests):
        return AgentReplyText(response="ok")

    async def stream(self, agent, session, requests):
        yield MessageStart(message_id="m-1")
        for token in ("Hel", "lo"):
            yield TextDelta(message_id="m-1", content=token)
        yield MessageEnd(message_id="m-1")

Every event reaches PostHook.on_stream_event(), but only TextDelta is projected into StreamChunk.delta — so a test that accumulates the reply must filter on chunk.delta is not None rather than slice by position. (delta is a model field, so the attribute always exists; it is None on every non-text frame. Key presence is the wire-format rule, which applies to the serialised SSE frames, not to the objects a test sees.)

A test double for a hook implements on_stream_event, and a double that returns a list ends the chain for that event — so a two-hook test asserting what the second hook saw is the way to pin that rule. See RecordingHook and HaltingHook in tests/test_runtime_stream_events.py.

Async Test Patterns

Use @pytest.mark.asyncio for async tests:

python
import pytest

@pytest.mark.asyncio
async def test_runtime_run():
    runtime = Runtime(InMemorySessionStore())
    agent = DummyAgent()
    runtime.register(agent)
    session = runtime.sessions().new("test-session")

    result = await runtime.run(agent, session, [AgentRequestText(prompt="hello")])
    assert result.response == "ok:hello"
Monkeypatching Config

Use monkeypatch to override AKConfig for tests:

python
def test_redis_session_store(monkeypatch):
    class FakeCfg:
        class session:
            type = "redis"
            cache = None
            class redis:
                url = "redis://localhost:6379"
                ttl = 60
                prefix = "ak:test:"

    monkeypatch.setattr("agentkernel.core.config.AKConfig.get", classmethod(lambda cls: FakeCfg))
    store = SessionStoreBuilder.build()
    assert isinstance(store, RedisSessionStore)
Session Context Tests

Test the async context manager pattern:

python
@pytest.mark.asyncio
async def test_session_context():
    session = Session("test-id")

    async with session:
        current = Session.current()
        assert current is session
        assert current.id == "test-id"

    # Outside context, no current session
    assert Session.current() is None
Testing Volatile vs Non-Volatile Caches
python
@pytest.mark.asyncio
async def test_volatile_cache_cleared():
    session = Session("test-id")

    async with session:
        session.get_volatile_cache().set("key", "value")
        assert session.get_volatile_cache().get("key") == "value"

    # Volatile cache is cleared after Runtime.run() completes
    # Non-volatile cache persists
Testing Hooks
python
@pytest.mark.asyncio
async def test_pre_hook_modifies_request():
    class TestPreHook(PreHook):
        async def on_run(self, session, agent, requests):
            for req in requests:
                if isinstance(req, AgentRequestText):
                    req.prompt = req.prompt.upper()
            return requests
        def name(self): return "test_hook"

    agent = DummyAgent()
    agent.pre_hooks.append(TestPreHook())
    # ... run through Runtime and verify modified input


@pytest.mark.asyncio
async def test_pre_hook_halts_execution():
    class BlockingHook(PreHook):
        async def on_run(self, session, agent, requests):
            return AgentReplyText(response="blocked", prompt="")
        def name(self): return "blocking_hook"

    # When a PreHook returns AgentReply, Runtime.run() returns it immediately
    # without calling the agent's runner

Built-in Test Framework

Agent Kernel provides a Test class (ak-py/src/agentkernel/test/) for integration testing. This framework is used in examples and can be used for testing deployed agents as well.

python
from agentkernel.test import Test

# In test files
@pytest_asyncio.fixture(scope="session", loop_scope="session")
async def test_client():
    test = Test("demo.py")       # Path to agent definition file
    await test.start()
    try:
        yield test
    finally:
        await test.stop()

@pytest.mark.order(1)
async def test_agent_response(test_client):
    await test_client.send("Who won the 1996 cricket world cup?")
    await test_client.expect(["Sri Lanka won the 1996 cricket world cup."])
Test Modes

Configured via test-config.yaml — a separate, un-nested file resolved from the cwd (or AK_TEST_CONFIG_PATH_OVERRIDE), loaded only when the test harness runs. It is not part of config.yaml; a leftover test: section there is ignored:

yaml
mode: score    # score | llm | fallback (default: fallback)
evaluator: deepeval   # built-in short name ('deepeval', 'opik' or 'jev'), or a dotted path to your own AKEvaluator subclass
llm:
  model: gpt-4o-mini
  provider: openai
  • score: Deterministic, offline scoring via the configured AKEvaluator's evaluate_by_score — the built-in DeepevalAKEvaluator uses Scorer.quasi_exact_match_score (normalised whole-string equality, no LLM call); the built-in OpikAKEvaluator uses Opik's LevenshteinRatio instead (graded fuzzy-similarity, not exact-match); the built-in JevAKEvaluator has no score mode (AKMetricNotSupported), so evaluator: jev needs mode: llm
  • llm: LLM-as-judge scoring via evaluate_by_llm — deepeval and opik use a GEval metric against the expected answer(s) (ground truth), while jev asks one TypeSafe Noul yes/no question (its probability is the score; comparison text is sent to api.typesafe.ai): DeepevalAKEvaluator via DeepEval's GEval and an LLMTestCase, OpikAKEvaluator via Opik's GEval and a single packed output string. Evaluation backends are pluggable (agentkernel.test.core.evaluator.AKEvaluator); see ak-py/src/agentkernel/test/test.py
  • fallback: Tries score first, falls back to llm if score fails

evaluator is resolved by Test._resolve_evaluator (agentkernel/test/test.py), the same bring-your-own-dotted-path pattern resolve_dotted/require_extra (core/util/factory.py) use for session stores, sandbox providers, and trace backends — a custom evaluator subclasses AKEvaluator (agentkernel.test.core.evaluator) and implements evaluate_by_score/ evaluate_by_llm, each returning an AKEvaluationResult. See examples/cli/custom-evaluator/ for a full worked example, and docs/specs/555-pluggable-test-evaluators/ for the design/spec behind this interface.

Test.compare/Test.expect also take return_metrics: bool = False — when True, they return the AKEvaluationResult instead of raising AssertionError on a failing comparison (other errors, e.g. AKEvaluationError/AKMetricNotSupported, still propagate).

Test.compare() for API Tests

For HTTP API tests, use Test.compare():

python
response = await http_client.send("What is 2+2?")
Test.compare(response, ["4", "The answer is 4"])

HTTP API Integration Tests

Pattern for testing deployed agents:

python
class APITestClient:
    def __init__(self, url):
        self.url = url
        self.session_id = str(uuid.uuid4())

    async def send(self, prompt, endpoint=""):
        payload = {
            "prompt": prompt,
            "session_id": self.session_id,
            "agent": "triage"
        }
        async with httpx.AsyncClient(timeout=30.0) as client:
            resp = await client.post(f"{self.url}{endpoint}", json=payload)
            resp.raise_for_status()
            return resp.json().get("result", "")

@pytest_asyncio.fixture(scope="session", loop_scope="session")
async def http_client():
    endpoint = os.getenv("AK_TEST_ENDPOINT")
    yield APITestClient(endpoint)

CI/CD Workflows

  • test.yaml: Triggers on pull requests, pushes to develop, and manual dispatch; has an update-lock-files job (dispatch-only) and a run-tests job that delegates to test-reusable.yaml
  • test-reusable.yaml: Reusable workflow (workflow_call) containing the actual test jobs, including the uv run pytest invocation and the transport-integration-tests job, which starts real NATS and Kafka containers from the transport examples' compose files (broker services only, up -d --wait nats|kafka) and runs tests/test_transport_contract_live.py against them on every PR. build-ak-py hands ak-py/dist to unit-tests and e2e-tests as a run-scoped artifact (actions/upload-artifact/download-artifact), deliberately not actions/cache: the workflow is called with an arbitrary checkout_ref from privileged contexts (workflow_dispatch on develop, pull_request_target for labelled fork PRs), where a cache write is a poisoning vector the default branch would later restore (CodeQL actions/cache-poisoning). Don't "optimise" it back into a cache
  • chart-test.yaml: The #495 Helm chart gates, path-filtered on the chart, the k8s example, and agentkernel/pipeline/: a lint job (ct lint via ak-deployment/ak-k8s/ci/ct.yaml, plus explicit helm template renders of every flavor values file and optional tier) and a kind-smoke matrix over the dev/baremetal/eks flavors, each installing the chart with the example images (the dev flavor through examples/k8s/openai-queue-mode/deploy/deploy.sh local, the example's own install script, so the user-facing path is what CI exercises against the checkout's chart) and driving one real chat request through NATS (baremetal/eks install the Gateway API CRDs and a live NACK controller so autoProvision: false verifies operator-reconciled objects; the CI overlays in ak-deployment/ak-k8s/ci/ swap only storage class and sizing). Skipped for fork PRs (needs OPENAI_API_KEY)
  • test-trusted-pr.yaml: Runs test-reusable.yaml with secrets for fork PRs that have been reviewed and labeled safe-to-test (pull_request_target)
  • test-github-app.yaml: Manual dispatch only; verifies the GitHub App secrets (APP_ID/APP_PRIVATE_KEY) are configured correctly
  • integration-test.yaml: "Nightly" (tier nightly) integration tests against deployed environments; scheduled weekly on Sundays at 5:30 PM UTC (cron: '30 17 * * 0'), plus manual dispatch
  • integration-test-weekly.yaml: Weekly integration tests against deployed environments (cron currently commented out; manual dispatch, with option to keep cloud resources on failure), plus two additional manual-dispatch-only jobs gated on inputs.provision_e2e_messaging: e2e-messaging-deploy (builds and Terraform-applies the e2e/app messaging harness to AWS ECS, then waits for the service to stabilize) and e2e-messaging-test (probes the deployed webhook, registers the Telegram webhook, then runs the e2e/tests pytest suite against Slack/Telegram/WhatsApp/Messenger/Instagram/Gmail). See e2e/README.md for the full harness design.
  • code-quality.yml: Runs linting checks (see code-quality skill)
  • pr-title-check.yaml: Fails the PR unless its title follows Conventional Commits (type: description / type(scope): description, types in CONTRIBUTING.md). Runs on pull_request open/edit/synchronize/reopen/ready-for-review so a result exists for every head SHA (required checks are recorded per SHA)
  • copilot-review-request.yaml: Requests a Copilot code review on open/reopen/ready-for-review using the COPILOT_REVIEW_PAT secret (pull_request_target, API-only, no checkout): an org-scoped fine-grained PAT (resource owner yaalalabs, this repo, Pull requests: Read and write) created by a licensed maintainer. It is separate from COPILOT_REQUEST_TOKEN because the account-level Copilot Requests permission only exists on user-owned tokens, which cannot hold org repository permissions. Exists because the develop ruleset's copilot_code_review rule only fires for authors who hold a Copilot license. Only collaborator-authored PRs are requested automatically: the script calls repos.checkCollaborator (needs only the Metadata read every fine-grained PAT carries) rather than trusting the payload's author_association, because each review is a premium request billed to the PAT owner. Also skips PRs where Copilot is already requested and bot-authored PRs. workflow_dispatch takes a PR number, bypasses the collaborator check, and is how a maintainer requests a review on an outside contributor's PR after reading it (or backfills one)
  • reviewed-label-reset.yaml: Removes the maintainer-applied Reviewed label whenever new commits are pushed (pull_request_target: synchronize, API-only, no checkout) so the PR reappears in the -label:Reviewed review queue

Every workflow declares an explicit least-privilege permissions block (CodeQL actions/missing-workflow-permissions); jobs that write repo content (commits, branches, PRs) do so with the GitHub App token minted in-job, not with GITHUB_TOKEN. Metadata-only writes such as the label removal in reviewed-label-reset.yaml use a job-scoped GITHUB_TOKEN (pull-requests: write) instead of minting an App token. When adding a workflow, start from permissions: {} (or contents: read if it checks out) and widen only where a step provably needs it. Workflows on pull_request_target (test-trusted-pr.yaml, copilot-review-request.yaml, reviewed-label-reset.yaml) must never check out or execute PR code unless gated behind a maintainer-applied label, as test-trusted-pr.yaml is.

scripts/update_examples_version.py keeps the agentkernel pin in sync across both examples/** and e2e/app (via --examples-dir e2e/app): publish.yaml's "Update e2e messaging harness with new version" step bumps the pin with --skip-lock right after a production publish (the just-published version isn't resolvable on PyPI yet, so the lock refresh is deferred), and test.yaml's "Update e2e messaging harness lock file" step later runs --force-lock --examples-dir e2e/app to regenerate e2e/app/uv.lock and commits it alongside examples/**/uv.lock. If you add a new pinned dependency consumer under e2e/app, make sure it's covered by this same version-bump/lock-refresh pair instead of drifting out of sync with the published package.

Both integration workflows restore the branch-built agentkernel wheel from the ak-py-${{ github.sha }} cache with fail-on-cache-miss: true — the job fails loudly instead of silently falling back to the published PyPI wheel if the build/cache step didn't run first. .github/scripts/run_single_test.py's test_aws_deployment() then runs ./build.sh local in the example directory (force-reinstalling the local wheel with --no-cache-dir before packaging/deploying) and invokes the test client with uv run --no-sync pytest ... so uv doesn't re-sync the venv from uv.lock and revert the local wheel back to the PyPI version. When adding a new example to integration-test-config.yaml, make sure its build.sh/deploy.sh local branch force-reinstalls agentkernel from ../../../ak-py/dist with --no-cache-dir, matching this pattern — otherwise the test can silently exercise a stale published version instead of the branch's code. The Kubernetes examples apply the same split to the chart: deploy/deploy.sh installs the published OCI chart at the pinned release version, deploy/deploy.sh local installs ak-deployment/ak-k8s/chart from the checkout, and both examples/sandbox/broker-nats/app_test.py and chart-test.yaml's dev-flavor smoke call the local form.

integration-test-weekly.yaml's AWS matrix jobs reuse the shared VPC/subnets/security groups created by the aws-serverless base deployment (examples/aws-serverless/openai, itself queue_mode = true with a request handler, agent runner, and response handler tier) instead of each creating its own: get-base-outputs runs .github/scripts/get_base_outputs.py --base-path examples/aws-serverless/openai --deploy-dir deploy to read vpc_id/private_subnet_ids plus the three per-tier *_security_group_id terraform outputs, and each matrix job's Deploy/Destroy steps forward them as run_single_test.py --vpc-id --private-subnet-ids --request-handler-security-group-id (every aws-serverless job — the request handler tier is also shared by the authorizer and WebSocket connection handler Lambdas) and additionally --agent-runner-security-group-id --response-handler-security-group-id (only scalable-openai/schedule-openai, the two other weekly-matrix examples with those tiers). These become TF_VAR_request_handler_security_group_id/TF_VAR_agent_runner_security_group_id/TF_VAR_response_handler_security_group_id, which the example's own deploy/variables.tf (default null) and deploy/main.tf (security_group_id = var.*_security_group_id) must declare and forward into the corresponding serverless_agents module object (request_handler/agent_runner/response_handler), per docs/specs/716-reuse-sg-in-integration-test-pipeline/. When adding a new aws-serverless example to the weekly matrix, wire its deploy/variables.tf/main.tf the same way for whichever tiers it has, or the job will keep creating (and paying for/leaking) its own SGs instead of reusing the base's.

integration-test-weekly.yaml's deploy/test/destroy are separate workflow steps (not one bash block) so each shows up as its own status: the Test step only runs if: steps.deploy.outcome == 'success', and Destroy runs if: !cancelled() && !(keep_resources_on_failure && (deploy or test failed)). .github/scripts/generate_test_matrix.py assigns each gcp-* matrix entry a deploy_stagger (90s apart) so parallel GCP jobs don't provision VPC connectors on the same network simultaneously; the workflow sleeps that many seconds before deploying. Known infra-flakiness mitigations to preserve when touching this area, since they were added specifically to fix recurring weekly e2e failures:

  • AWS containerized examples' deploy.sh call wait_for_ecs_stable after terraform apply (reads region/prefix from terraform.tfvars, then aws ecs wait services-stable) so the test doesn't hit the app before the ECS service has finished rolling out.
  • run_single_test.py's destroy_aws_resources pre-deletes Lambda functions on the example's security groups and runs a background thread that periodically deletes their now-detached ENIs, so terraform destroy isn't blocked waiting ~20 min for AWS to release Hyperplane ENIs.
  • run_single_test.py's deploy_gcp_resources/destroy_gcp_resources retry the deploy up to 3 times and sweep ERROR-state VPC Access Connectors between attempts (sweep_gcp_error_connectors, matched by network name from terraform.tfvars).
  • Azure serverless/containerized deploys accept an AK_PRE_DEPLOY_AUTH_CMD env var (set in the workflows to refresh the OIDC-based az login) and re-run it before each long-running az step in linux_function.tf's deploy_function_code, since the OIDC token can expire mid-deploy.

If a weekly/nightly integration run fails with a symptom matching one of these (ECS "still rolling out", GCP VPC_ACCESS_CONNECTOR_ERROR, AWS destroy timing out on ENIs, or Azure CLI auth expiring), look here before re-adding ad-hoc sleeps or retries.

Best Practices

  1. Use ordered tests (@pytest.mark.order(n)) when testing conversational flows where follow-up questions depend on prior context
  2. Use session-scoped fixtures for test clients that are expensive to create
  3. Mock external services (LLM APIs, cloud services) in unit tests — only hit real APIs in integration tests
  4. Test both success and failure paths — especially for hooks and guardrails
  5. Use DummyAgent/DummyRunner to isolate the component under test from framework-specific behavior
  6. Test session persistence — verify that state survives across multiple Runtime.run() calls within the same session

© yaalalabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/ak-dev-testing-conventions of yaalalabs/agent-kernel.

Open the folder on GitHubat commit e03a602

Compare with similar skills

Ak Dev Testing Conventions next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ak Dev Testing Conventions compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ak Dev Testing Conventions this skillyaalalabs/agent-kernel191—~14kAutomated safety check: PassApache-2.0
Blue Teamgaasher/Agent-Loop-Skills174—~3.6kAutomated safety check: PassMIT
Simple Modern Uvjlevy/simple-modern-uv301—~1.9kAutomated safety check: PassMIT
Gating Deid Leakagemaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
Run Testsmicrosoft/testfx1k—~4.2kAutomated safety check: PassMIT
Megatron Core Testing GuideNVIDIA/Megatron-LM18k1 repos~2.1kAutomated safety check: PassApache-2.0

Similar skills

  • Blue Team

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…

    174 GitHub stars~3.6k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Simple Modern Uv

    jlevy/simple-modern-uv

    Start, selectively modernize, fully migrate, or update Python projects using simple-modern-uv practices: uv, ruff, BasedPyright, pytest, GitHub Actions CI, and tag-driven PyPI publishing.

    301 GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Gating Deid Leakage

    maziyarpanahi/openmed

    Add a CI gate that fails the build when an OpenMed de-identification model's recall on a held-out PHI set drops below threshold or any critical identifier leaks.

    5.5k GitHub stars~1.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Run Tests

    microsoft/testfx

    Official

    For dotnet test: figures out which test platform (VSTest vs Microsoft.Testing.Platform) a project uses from Directory.Build.props, global.json, and .csproj, then picks the matching command syntax.

    1k GitHub stars~4.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Megatron Core Testing Guide

    NVIDIA/Megatron-LM

    Official

    Guide to the Megatron-LM test system: layout, recipe YAML, running and adding unit and functional tests, golden values, marker filters and CI parity.

    18k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check passed
  • Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…

    2.8k GitHub stars~2.2k tokensUpdated today
    Testing & QAAuto-check passed

More from yaalalabs/agent-kernel

All 23 skills in this repo
  • Ak Dev Code Quality

    yaalalabs/agent-kernel

    Code quality standards, formatting, Python style rules (classes over script-style functions, configuration-field rules), commit conventions, and PR workflow for Agent Kernel development.

    191 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Ak Dev New Evaluator Provider

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).

    191 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ak Dev New Guardrail Provider

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new guardrail provider to Agent Kernel.

    191 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Step-by-step guide for adding a new knowledge base backend to Agent Kernel.

    191 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Ak Dev New Messaging Integration

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new messaging platform integration to Agent Kernel.

    191 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Ak Dev New Multimodal Storage

    yaalalabs/agent-kernel

    Step-by-step guide for adding a new multimodal attachment storage backend to Agent Kernel.

    191 GitHub stars~4.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Ak Dev Testing Conventions

What does Ak Dev Testing Conventions do?

Testing conventions, patterns, and automation for Agent Kernel development. Ak Dev Testing Conventions is an agent skill from yaalalabs/agent-kernel. Testing conventions, patterns, and automation for Agent Kernel development.

When should I use Ak Dev Testing Conventions?

Ak Dev Testing Conventions fits situations like: writing tests for new features; debugging test failures; understanding the test infrastructure.

How do I install Ak Dev Testing Conventions in Claude Code?

Run `npx skills add yaalalabs/agent-kernel --skill ak-dev-testing-conventions -a claude-code`. Or copy the skill folder (.agents/skills/ak-dev-testing-conventions in yaalalabs/agent-kernel) into .claude/skills/ak-dev-testing-conventions in your project. Claude Code loads it when a task matches its description.

How do I install Ak Dev Testing Conventions in Codex?

Run `npx skills add yaalalabs/agent-kernel --skill ak-dev-testing-conventions -a codex`. Or copy the skill folder (.agents/skills/ak-dev-testing-conventions in yaalalabs/agent-kernel) into .agents/skills/ak-dev-testing-conventions in your project. Codex loads it when a task matches its description.

Can I use Ak Dev Testing Conventions in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yaalalabs/agent-kernel --skill ak-dev-testing-conventions -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ak-dev-testing-conventions, .gemini/skills/ak-dev-testing-conventions, .github/skills/ak-dev-testing-conventions and .opencode/skills/ak-dev-testing-conventions in your project.

What does Ak Dev Testing Conventions need to run?

Going by SKILL.md and its folder, Ak Dev Testing Conventions needs the command-line tools its instructions call (uv, terraform, helm, aws and az) and credentials named GITHUB_TOKEN, OPENAI_API_KEY, APP_PRIVATE_KEY and COPILOT_REQUEST_TOKEN.

Does Ak Dev Testing Conventions access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ak Dev Testing Conventions safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ak Dev Testing Conventions use?

Ak Dev Testing Conventions is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ak Dev Testing Conventions use?

About 14k tokens (SKILL.md is roughly 55k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ak Dev Testing Conventions?

Skills that share tags, products or a category with Ak Dev Testing Conventions: Blue Team (gaasher/Agent-Loop-Skills, 174 stars), Simple Modern Uv (jlevy/simple-modern-uv, 301 stars), Gating Deid Leakage (maziyarpanahi/openmed, 5.5k stars) and Run Tests (microsoft/testfx, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ak Dev Testing Conventions?

yaalalabs (a GitHub organization) maintains it in yaalalabs/agent-kernel, which has 191 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 8, 2026.

Source: yaalalabs/agent-kernel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.