Smt E2E Dataflow Debugging
GoogleCloudPlatform/DataflowTemplates
Debugs logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and comparing source (e.g.
Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU).
$ npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google-ai-edge/LiteRT debugging-litert-gpu-accuracy --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google-ai-edge/LiteRT.git skills-src && mkdir -p .claude/skills && cp -r skills-src/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy .claude/skills/debugging-litert-gpu-accuracy && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "debugging-litert-gpu-accuracy" agent skill from https://github.com/google-ai-edge/LiteRT/tree/main/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy into .claude/skills/debugging-litert-gpu-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-litert-gpu-accuracy", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google-ai-edge/LiteRT/tree/main/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google-ai-edge/LiteRT debugging-litert-gpu-accuracy --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-ai-edge/LiteRT.git skills-src && mkdir -p .agents/skills && cp -r skills-src/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy .agents/skills/debugging-litert-gpu-accuracy && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "debugging-litert-gpu-accuracy" agent skill from https://github.com/google-ai-edge/LiteRT/tree/main/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy into .agents/skills/debugging-litert-gpu-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-litert-gpu-accuracy", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google-ai-edge/LiteRT debugging-litert-gpu-accuracy --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-ai-edge/LiteRT.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy .cursor/skills/debugging-litert-gpu-accuracy && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "debugging-litert-gpu-accuracy" agent skill from https://github.com/google-ai-edge/LiteRT/tree/main/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy into .cursor/skills/debugging-litert-gpu-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-litert-gpu-accuracy", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google-ai-edge/LiteRT.git --path google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google-ai-edge/LiteRT debugging-litert-gpu-accuracy --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-ai-edge/LiteRT.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy .gemini/skills/debugging-litert-gpu-accuracy && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "debugging-litert-gpu-accuracy" agent skill from https://github.com/google-ai-edge/LiteRT/tree/main/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy into .gemini/skills/debugging-litert-gpu-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-litert-gpu-accuracy", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google-ai-edge/LiteRT debugging-litert-gpu-accuracyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google-ai-edge/LiteRT.git skills-src && mkdir -p .github/skills && cp -r skills-src/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy .github/skills/debugging-litert-gpu-accuracy && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "debugging-litert-gpu-accuracy" agent skill from https://github.com/google-ai-edge/LiteRT/tree/main/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy into .github/skills/debugging-litert-gpu-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-litert-gpu-accuracy", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google-ai-edge/LiteRT debugging-litert-gpu-accuracy --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google-ai-edge/LiteRT.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy .opencode/skills/debugging-litert-gpu-accuracy && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "debugging-litert-gpu-accuracy" agent skill from https://github.com/google-ai-edge/LiteRT/tree/main/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy into .opencode/skills/debugging-litert-gpu-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-litert-gpu-accuracy", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
debugging-litert-gpu-accuracyDiagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU).
Debugging Litert GPU Accuracy is an agent skill from google-ai-edge/LiteRT. Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU). Use when a model produces correct output on CPU but generates garbage, NaNs, or wrong outputs on GPU; when comparing CPU vs GPU execution op-by-op or state-buffer-by-state-buffer; or when debugging composite ops (cacheupdate, runtimebmm, sdpa), ring-buffer sliding-window KV caches, or TensorStorageType (BUFFER vs…
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/contributing.md` and `references/mldrift_shader_probes.md`).
It sits in Development, covering Debugging and End-to-end testing. It works with Android. The repository describes itself as: LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and…. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 320bfd5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
adbFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Debugging Litert GPU Accuracy loads about 3.3k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 182 tokens; SKILL.md has 987 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google-ai-edge/LiteRT at commit 320bfd5, republished under its Apache-2.0 licence (© google-ai-edge). 987 words, ~3,250 tokens.
.claude/skills/debugging-litert-gpu-accuracy/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Use this 6-phase funnel to isolate why a LiteRT-LM model produces correct
output on CPU (--backend=cpu) but generates garbage, NaNs, or empty/corrupted
state on Android GPU (--backend=gpu, OpenCL or WebGPU).
Phase 1: Unpack & diff model metadata (.litertlm) against a GPU-good model
│
Phase 2: On-device repro matrix (CPU vs OpenCL vs WebGPU + prompt sweep)
│
Phase 3: Op-by-op accuracy_debugger on extracted TFLite subgraph
│ ├─ Op diverges in isolation ──► Fix op kernel / precision
│ └─ All ops pass (maxdiff≈0) ──► Composite / graph-lowering / state bug
▼
Phase 4: Host-side CPU vs GPU state-buffer dump after prefill_128 & decode
│ (Pinpoints exact corrupted KV cache layers: e.g. local ring vs global)
▼
Phase 5: ML Drift descriptor dump & in-shader sentinel readback probes
│ (Bisects runtime param chain & TensorStorageType mismatches)
▼
Phase 6: Fix in ml_drift / delegate, unit-test GpuModel invariants, & validateDelete the on-device ML Drift program cache before every GPU run after editing any kernel, parser, or selector code. Otherwise the device silently reuses the cached compiled OpenCL/WebGPU program and your C++ edits never execute:
adb shell "rm -f /data/local/tmp/{model_dir}/*mldrift_program_cache.bin"Keep *_mldrift_weight_cache.bin and *.xnnpack_cache so weight conversion
stays fast.
Injecting device environment variables: run_litert_lm.sh does not
forward custom host env vars or provide a build-only flag. Run
run_litert_lm.sh --skip_model_push once to build and push
libLiteRtOpenClAccelerator.so and litert_lm_advanced_main, then invoke
adb shell directly with your probe env vars:
OUT=/data/local/tmp/{model_dir}
adb shell "rm -f ${OUT}/*mldrift_program_cache.bin"
adb shell "ADSP_LIBRARY_PATH=${OUT} LD_LIBRARY_PATH=${OUT} \
LITERT_LM_DEBUG_DUMP_STATE=1 MLDRIFT_CACHE_PROBE=1 \
${OUT}/litert_lm_advanced_main \
--model_path=${OUT}/model.litertlm --backend=gpu \
--async=false --clear_kv_cache_before_prefill=true \
--convert_weights_on_gpu=true --minloglevel=0 \
--input_prompt=\"hi\""Layering check in third_party/ml_drift and delegate/composite:
Adding #include "third_party/absl/log/absl_log.h" to kernel files fails
Bazel/Blaze layering_check. Use #include <cstdio> + fprintf(stderr, ...) and #include <cstdlib> + std::getenv(...) for temporary probes.
Unpack the failing .litertlm bundle and a known-good GPU model to compare
their ExecutorMetadataProto, LlmMetadataProto, and embedded .tflite
subgraphs:
/google/bin/releases/arca9-local-blaze-cli/blaze-for-agents run \
//third_party/odml/litert_lm/runtime/util:litertlm_builder_cli -- \
unpack --input_file=/tmp/dbg/failing.litertlm --output_dir=/tmp/dbg/unpackedCheck ExecutorMetadataProto.pbtext and LlmMetadataProto.pbtext for:
TYPE_LOCAL_*_CACHE
(sliding-window / ring buffer, e.g. minimum_sequence_length: 1024,
maximum_sequence_length: 1280) with global caches
(maximum_sequence_length: 32771)?attention_mask_settings: Presence of sliding_window_size triggers
the gpu_optimized_single_buffer_cache_ ring-buffer path in LiteRT-LM and
ring_buffer_sdpa in the exported graph.Run the repro prompt across backends and configurations to eliminate whole subsystems up front:
CPU vs OpenCL GPU vs WebGPU:
# OpenCL GPU
./third_party/odml/litert_lm/runtime/engine/run_litert_lm.sh \
--models={model_name} --target_os=android --backend=gpu \
--skip_model_push --input_prompt="{prompt}"
# WebGPU (--use_webgpu)
./third_party/odml/litert_lm/runtime/engine/run_litert_lm.sh \
--models={model_name} --target_os=android --backend=gpu --use_webgpu \
--skip_model_push --input_prompt="{prompt}"ml_drift/common / delegate/composite graph
lowering (such as LiteRtOpSelector), not an OpenCL-specific driver
or shader bug.Prompt-length sweep: Test a minimal prompt ("hi" — note that system
prompt wrapping still yields ~14 tokens) vs >128 tokens (multi-chunk
prefill) vs >sliding_window_size. If even a 14-token prompt fails on the
very first generated token, wrap-around and multi-chunk boundary math are
ruled out.
accuracy_debugger (and Know Its Blind Spots)Run accuracy_debugger on the unpacked
Section3_TFLiteModel_tf_lite_prefill_decode.tflite (prefill_128 subgraph) to
compare GPU vs CPU reference op-by-op:
./litert/tools/accuracy_debugger/google/debug_accuracy_qc.sh \
--model_path=/tmp/dbg/unpacked/Section3_TFLiteModel_tf_lite_prefill_decode.tflite \
--signature_name=prefill_128 \
--num_ops=400Inspect the generated CSV (Index, Op Code, Tensor Name, Max Diff, MSE, Cos Sim, SNR, PSNR, NaN, Min Ref, Max Ref, Min Accel, Max Accel, Status):
Cos Sim / high Max Diff / NaN: You have
an isolated op numerical/precision bug.accuracy_debugger verifies each op in
isolation and reports ACCEL_COMPILE_FAILED on
shlo.composite(odml.runtime_bmm) and shlo.composite(odml.cache_update).
Consequently, an entire parameter computation chain (e.g. SLICE →
FLOOR_MOD → CONCAT) can pass with maxdiff = 0 in accuracy_debugger
while still producing garbage in the full compiled graph due to cross-op
storage-type / producer-rewiring bugs triggered only when composite ops
are lowered alongside standard ops.prefill_128 & decodeWhen accuracy_debugger shows standard ops are healthy, instrument
google3/third_party/odml/litert_lm/runtime/executor/llm_litert_compiled_model_executor.cc
(see references/mldrift_shader_probes.md
for the drop-in snippet) to dump n, sum, sumabs, min, max, nan, and
the first 6 elements (head=[...]) of every input and output buffer in
BindTensorsAndRunPrefill and BindTensorsAndRunDecode under --async=false.
Run once with --backend=cpu and once with --backend=gpu on "hi" and diff:
param_tensor (prefill_in): Verify the 7-channel int32
runtime parameter tensor [start, end, end, ...] matches on host.kv_cache_k_{layer} / kv_cache_v_{layer} (prefill_out):sumabs > 0) while local/sliding-window layers (e.g. layers 0, 1,
2, ...) have sumabs = 0 (min=0, max=0), the ring-buffer branch of
add_values_to_cache_kernel.cc is early-exiting without writing.Understand the parameter contract before probing:
| Slot | Meaning | Consumer |
|---|---|---|
0 | token_index_offset (start or start % S) | add_values_to_cache (both branches) |
1 | active_tokens (end or min(end, S)) | add_values_to_cache non-ring bound |
2 | kActiveTokensAlignedIndex | runtime_batched_matmul (src_end_ch_index / dst_end_ch_index) & sdpa_transposed |
3 | Ring update_length (end - start) | add_values_to_cache ring branch only (if (X >= update_length) return;) |
Use two targeted probes (full code in references/mldrift_shader_probes.md):
MLDRIFT_CACHE_PROBE): In
ml_drift/delegate/composite/add_values_to_cache_kernel.cc,
bypass the X >= update_length gate and overwrite final_value_k with
float4(params.Read(0..3)) and final_value_v with
float4(params.Read(4..7)). The host [DBG] dump of
kv_cache_k_0 head=[...] and kv_cache_v_0 head=[...] then directly prints
the exact int32 values the GPU shader read from args.params!OperationDef / TensorStorageType dump
(MLDRIFT_SLICE_PROBE): In
google3/third_party/ml_drift/common/kernels/strided_slice.cc (inside
GetStridedSliceCode), print starts.c, ends.c, src.Channels(), and
src.GetStorageType() (1 = BUFFER, 3 = TEXTURE_2D) for small int32
tensors (c <= 16).b/565413009)If MLDRIFT_CACHE_PROBE shows params[0..2] are valid (0, 14, 14) while
params[3..6] are garbage (±inf in fp16), and [SLICEDBG] shows:
starts.c=0 ends.c=3 | src c=7 storage=1 (BUFFER) -> valid
starts.c=3 ends.c=7 | src c=7 storage=3 (TEXTURE_2D) -> GARBAGE (unproduced tensor!)Look at LiteRtOpSelector::ParamTensorToBuffer in
ml_drift/delegate/composite/litert_op_selector.cc:
add_values_to_cache, runtime_batched_matmul,
sdpa_transposed) declare params via AddSrcBuffer (BUFFER storage),
whereas standard ops like StridedSlice (rest = param'[3:7]) consume the
original tensor via AddSrcTensor (TEXTURE_2D) without consulting
replaced_tensors_.model_builder->UpdateOutputTensor(param_tensor, new_param_tensor.id)
re-points the producer node's output from param_id to
new_param_tensor.id, leaving param_id with no producer in the
GpuModel so any non-buffer consumer reads uninitialized GPU memory.model_builder->Copy(param_tensor, new_param_tensor)
instead of UpdateOutputTensor. When param_tensor has no other consumers,
LinkNodes() in google3/third_party/ml_drift/common/merge_nodes.cc
automatically fuses the single-consumer copy back into the producer.GpuModel Invariant & Validate End-to-EndRevert all temporary probes (hg revert on kernel/executor files)
before running final validation.
Write a hermetic cc_test (see
ml_drift/delegate/composite/litert_op_selector_test.cc)
that builds a GpuModel via GpuModelBuilder + LiteRtOpSelector and
asserts that every in_id of every GpuNode in gpu_model.nodes is
present in defined_tensors (graph inputs ∪ const tensors ∪ node
outputs):
/google/bin/releases/arca9-local-blaze-cli/blaze-for-agents test \
//ml_drift/delegate/composite:litert_op_selector_testValidate on device across:
--backend=gpu)--backend=gpu --use_webgpu)--models=kanana_i4_107_kv_1024 --backend=gpu)To contribute or modify this skill or its references, follow the contribution guidelines before making changes.
© google-ai-edge, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy of google-ai-edge/LiteRT.
Open the folder on GitHubat commit 320bfd5
Debugging Litert GPU Accuracy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Debugging Litert GPU Accuracy this skillgoogle-ai-edge/LiteRT | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Smt E2E Dataflow DebuggingGoogleCloudPlatform/DataflowTemplates | 1.3k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Debugging MarchatCod-e-Codes/marchat | 137 | — | ~668 | Automated safety check: Notes | MIT | |
| Alefxberg-io/alef | 100 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Try Fix Alternative Approachdotnet/maui | 23k | — | ~8.4k | Automated safety check: Pass | MIT | |
| CanvasBitterbot-AI/bitterbot-desktop | 2.5k | — | ~1.4k | Automated safety check: Pass | MIT |
GoogleCloudPlatform/DataflowTemplates
Debugs logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and comparing source (e.g.
Cod-e-Codes/marchat
Diagnoses marchat client and server issues using -doctor, env configuration, and logs.
xberg-io/alef
Use Alef correctly for Rust-to-polyglot binding generation. An agent skill from xberg-io/alef.
dotnet/maui
Attempts one alternative fix for a bug, runs the given test command against it and reports what happened, always differing from existing PR fixes.
Bitterbot-AI/bitterbot-desktop
Display and control HTML content on connected Bitterbot nodes (Mac, iOS, Android) via the canvas host server.
callstackincubator/agent-skills
Systematically explore and test a mobile app on iOS/Android with agent-device to find bugs, UX issues, and other problems.
google-ai-edge/LiteRT
Authors, registers, and tests new operator test generators for the LiteRT Accelerator Test Suite (ATS) in litert/test/generators/ and litert/ats/.
Works with
Categories
Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU). Debugging Litert GPU Accuracy is an agent skill from google-ai-edge/LiteRT. Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU).
Debugging Litert GPU Accuracy fits situations like: A model produces correct output on CPU but generates garbage; wrong outputs on GPU; comparing CPU vs GPU execution op-by-op; state-buffer-by-state-buffer.
Run `npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a claude-code`. Or copy the skill folder (google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy in google-ai-edge/LiteRT) into .claude/skills/debugging-litert-gpu-accuracy in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a codex`. Or copy the skill folder (google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy in google-ai-edge/LiteRT) into .agents/skills/debugging-litert-gpu-accuracy in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging-litert-gpu-accuracy, .gemini/skills/debugging-litert-gpu-accuracy, .github/skills/debugging-litert-gpu-accuracy and .opencode/skills/debugging-litert-gpu-accuracy in your project.
Going by SKILL.md and its folder, Debugging Litert GPU Accuracy needs the command-line tools its instructions call (adb).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Debugging Litert GPU Accuracy is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Debugging Litert GPU Accuracy: Smt E2E Dataflow Debugging (GoogleCloudPlatform/DataflowTemplates, 1.3k stars), Debugging Marchat (Cod-e-Codes/marchat, 137 stars), Alef (xberg-io/alef, 100 stars) and Try Fix Alternative Approach (dotnet/maui, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google-ai-edge (a GitHub organization) maintains it in google-ai-edge/LiteRT, which has 3,475 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 10, 2026.
Source: google-ai-edge/LiteRT on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.