Agent skill

Debugging Litert GPU Accuracy

by google-ai-edge in google-ai-edge/LiteRT

Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU).

Apache-2.0Auto-check passedDevelopment

Install Debugging Litert GPU Accuracy

skills CLI
$ npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-ai-edge/LiteRT debugging-litert-gpu-accuracy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-ai-edge/LiteRT.git skills-src && mkdir -p .claude/skills && cp -r skills-src/google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy .claude/skills/debugging-litert-gpu-accuracy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debugging-litert-gpu-accuracy
GitHub stars
3.5k
Token cost
~3.3k tokens
SKILL.md length
987 words
Files
3 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU).

  • Works in 6 steps: Unpack & Diff Model Metadata Against a… → Reproduce & Narrow the Failure Surface… → Run Op-by-Op accuracy_debugger (and Know… → …
  • A model produces correct output on CPU but generates garbage
  • SKILL.md covers Critical Environment Gotchas…, Phase 1: Unpack & Diff Model…, Phase 2: Reproduce & Narrow… and Phase 3: Run Op-by-Op…, plus 5 more sections
  • Calls adb

What it does

Debugging Litert GPU Accuracy is an agent skill from google-ai-edge/LiteRT. Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU). Use when a model produces correct output on CPU but generates garbage, NaNs, or wrong outputs on GPU; when comparing CPU vs GPU execution op-by-op or state-buffer-by-state-buffer; or when debugging composite ops (cacheupdate, runtimebmm, sdpa), ring-buffer sliding-window KV caches, or TensorStorageType (BUFFER vs…

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/contributing.md` and `references/mldrift_shader_probes.md`).

It sits in Development, covering Debugging and End-to-end testing. It works with Android. The repository describes itself as: LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and…. The licence is Apache-2.0.

When your agent uses it

  • A model produces correct output on CPU but generates garbage
  • Wrong outputs on GPU
  • Comparing CPU vs GPU execution op-by-op
  • State-buffer-by-state-buffer

Example prompts

  • “Use the debugging-litert-gpu-accuracy skill to diagnose and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML…”
  • “/debugging-litert-gpu-accuracy”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Unpack & Diff Model Metadata Against a Known-Good Model
  2. Reproduce & Narrow the Failure Surface on Device
  3. Run Op-by-Op accuracy_debugger (and Know Its Blind Spots)
  4. Diff CPU vs GPU State Buffers After prefill_128 & decode
  5. Bisect with ML Drift Descriptor & Shader Readback Probes
  6. Unit-Test the GpuModel Invariant & Validate End-to-End

What it can do on your machine

Read from SKILL.md and the folder at commit 320bfd5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • adb

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debugging Litert GPU Accuracy loads about 3.3k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 182 tokens; SKILL.md has 987 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~182
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google-ai-edge/LiteRT at commit 320bfd5, republished under its Apache-2.0 licence (© google-ai-edge). 987 words, ~3,250 tokens.

Download SKILL.mdSave it as .claude/skills/debugging-litert-gpu-accuracy/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
debugging-litert-gpu-accuracy
description
Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU). Use when a model produces correct output on CPU but generates garbage, NaNs, or wrong outputs on GPU; when comparing CPU vs GPU execution op-by-op or state-buffer-by-state-buffer; or when debugging composite ops (cache_update, runtime_bmm, sdpa), ring-buffer sliding-window KV caches, or TensorStorageType (BUFFER vs TEXTURE_2D) mismatches in ml_drift. Don't use for NPU/QNN dispatch issues, CPU-only model conversion failures, or adding standard E2E device test targets (use adding-litert-lm-e2e-tests).

Debugging LiteRT / LiteRT-LM On-Device GPU Accuracy & ML Drift Bugs

Use this 6-phase funnel to isolate why a LiteRT-LM model produces correct output on CPU (--backend=cpu) but generates garbage, NaNs, or empty/corrupted state on Android GPU (--backend=gpu, OpenCL or WebGPU).

Phase 1: Unpack & diff model metadata (.litertlm) against a GPU-good model
   │
Phase 2: On-device repro matrix (CPU vs OpenCL vs WebGPU + prompt sweep)
   │
Phase 3: Op-by-op accuracy_debugger on extracted TFLite subgraph
   │     ├─ Op diverges in isolation ──► Fix op kernel / precision
   │     └─ All ops pass (maxdiff≈0) ──► Composite / graph-lowering / state bug
   ▼
Phase 4: Host-side CPU vs GPU state-buffer dump after prefill_128 & decode
   │     (Pinpoints exact corrupted KV cache layers: e.g. local ring vs global)
   ▼
Phase 5: ML Drift descriptor dump & in-shader sentinel readback probes
   │     (Bisects runtime param chain & TensorStorageType mismatches)
   ▼
Phase 6: Fix in ml_drift / delegate, unit-test GpuModel invariants, & validate

Critical Environment Gotchas (Read First)

  1. Delete the on-device ML Drift program cache before every GPU run after editing any kernel, parser, or selector code. Otherwise the device silently reuses the cached compiled OpenCL/WebGPU program and your C++ edits never execute:

    bash
    adb shell "rm -f /data/local/tmp/{model_dir}/*mldrift_program_cache.bin"

    Keep *_mldrift_weight_cache.bin and *.xnnpack_cache so weight conversion stays fast.

  2. Injecting device environment variables: run_litert_lm.sh does not forward custom host env vars or provide a build-only flag. Run run_litert_lm.sh --skip_model_push once to build and push libLiteRtOpenClAccelerator.so and litert_lm_advanced_main, then invoke adb shell directly with your probe env vars:

    bash
    OUT=/data/local/tmp/{model_dir}
    adb shell "rm -f ${OUT}/*mldrift_program_cache.bin"
    adb shell "ADSP_LIBRARY_PATH=${OUT} LD_LIBRARY_PATH=${OUT} \
      LITERT_LM_DEBUG_DUMP_STATE=1 MLDRIFT_CACHE_PROBE=1 \
      ${OUT}/litert_lm_advanced_main \
      --model_path=${OUT}/model.litertlm --backend=gpu \
      --async=false --clear_kv_cache_before_prefill=true \
      --convert_weights_on_gpu=true --minloglevel=0 \
      --input_prompt=\"hi\""
  3. Layering check in third_party/ml_drift and delegate/composite: Adding #include "third_party/absl/log/absl_log.h" to kernel files fails Bazel/Blaze layering_check. Use #include <cstdio> + fprintf(stderr, ...) and #include <cstdlib> + std::getenv(...) for temporary probes.


Phase 1: Unpack & Diff Model Metadata Against a Known-Good Model

Unpack the failing .litertlm bundle and a known-good GPU model to compare their ExecutorMetadataProto, LlmMetadataProto, and embedded .tflite subgraphs:

bash
/google/bin/releases/arca9-local-blaze-cli/blaze-for-agents run \
  //third_party/odml/litert_lm/runtime/util:litertlm_builder_cli -- \
  unpack --input_file=/tmp/dbg/failing.litertlm --output_dir=/tmp/dbg/unpacked

Check ExecutorMetadataProto.pbtext and LlmMetadataProto.pbtext for:

  • Hybrid KV cache topology: Does the model mix TYPE_LOCAL_*_CACHE (sliding-window / ring buffer, e.g. minimum_sequence_length: 1024, maximum_sequence_length: 1280) with global caches (maximum_sequence_length: 32771)?
  • attention_mask_settings: Presence of sliding_window_size triggers the gpu_optimized_single_buffer_cache_ ring-buffer path in LiteRT-LM and ring_buffer_sdpa in the exported graph.

Phase 2: Reproduce & Narrow the Failure Surface on Device

Run the repro prompt across backends and configurations to eliminate whole subsystems up front:

  1. CPU vs OpenCL GPU vs WebGPU:

    bash
    # OpenCL GPU
    ./third_party/odml/litert_lm/runtime/engine/run_litert_lm.sh \
      --models={model_name} --target_os=android --backend=gpu \
      --skip_model_push --input_prompt="{prompt}"
    
    # WebGPU (--use_webgpu)
    ./third_party/odml/litert_lm/runtime/engine/run_litert_lm.sh \
      --models={model_name} --target_os=android --backend=gpu --use_webgpu \
      --skip_model_push --input_prompt="{prompt}"
    • If both OpenCL and WebGPU fail identically, the bug is in backend-agnostic code: the exported graph, LiteRT-LM host tensor filling, or shared ml_drift/common / delegate/composite graph lowering (such as LiteRtOpSelector), not an OpenCL-specific driver or shader bug.
  2. Prompt-length sweep: Test a minimal prompt ("hi" — note that system prompt wrapping still yields ~14 tokens) vs >128 tokens (multi-chunk prefill) vs >sliding_window_size. If even a 14-token prompt fails on the very first generated token, wrap-around and multi-chunk boundary math are ruled out.


Phase 3: Run Op-by-Op accuracy_debugger (and Know Its Blind Spots)

Run accuracy_debugger on the unpacked Section3_TFLiteModel_tf_lite_prefill_decode.tflite (prefill_128 subgraph) to compare GPU vs CPU reference op-by-op:

bash
./litert/tools/accuracy_debugger/google/debug_accuracy_qc.sh \
  --model_path=/tmp/dbg/unpacked/Section3_TFLiteModel_tf_lite_prefill_decode.tflite \
  --signature_name=prefill_128 \
  --num_ops=400

Inspect the generated CSV (Index, Op Code, Tensor Name, Max Diff, MSE, Cos Sim, SNR, PSNR, NaN, Min Ref, Max Ref, Min Accel, Max Accel, Status):

  • If a standard op shows low Cos Sim / high Max Diff / NaN: You have an isolated op numerical/precision bug.
  • Blind spot to remember: accuracy_debugger verifies each op in isolation and reports ACCEL_COMPILE_FAILED on shlo.composite(odml.runtime_bmm) and shlo.composite(odml.cache_update). Consequently, an entire parameter computation chain (e.g. SLICE → FLOOR_MOD → CONCAT) can pass with maxdiff = 0 in accuracy_debugger while still producing garbage in the full compiled graph due to cross-op storage-type / producer-rewiring bugs triggered only when composite ops are lowered alongside standard ops.

Phase 4: Diff CPU vs GPU State Buffers After prefill_128 & decode

When accuracy_debugger shows standard ops are healthy, instrument google3/third_party/odml/litert_lm/runtime/executor/llm_litert_compiled_model_executor.cc (see references/mldrift_shader_probes.md for the drop-in snippet) to dump n, sum, sumabs, min, max, nan, and the first 6 elements (head=[...]) of every input and output buffer in BindTensorsAndRunPrefill and BindTensorsAndRunDecode under --async=false.

Run once with --backend=cpu and once with --backend=gpu on "hi" and diff:

  • Check param_tensor (prefill_in): Verify the 7-channel int32 runtime parameter tensor [start, end, end, ...] matches on host.
  • Check kv_cache_k_{layer} / kv_cache_v_{layer} (prefill_out):
    • If global layers (e.g. layers 3, 7, 11, ...) match CPU (sumabs > 0) while local/sliding-window layers (e.g. layers 0, 1, 2, ...) have sumabs = 0 (min=0, max=0), the ring-buffer branch of add_values_to_cache_kernel.cc is early-exiting without writing.

Show full SKILL.md (357 more words)Show less

Phase 5: Bisect with ML Drift Descriptor & Shader Readback Probes

Understand the parameter contract before probing:

SlotMeaningConsumer
0token_index_offset (start or start % S)add_values_to_cache (both branches)
1active_tokens (end or min(end, S))add_values_to_cache non-ring bound
2kActiveTokensAlignedIndexruntime_batched_matmul (src_end_ch_index / dst_end_ch_index) & sdpa_transposed
3Ring update_length (end - start)add_values_to_cache ring branch only (if (X >= update_length) return;)

Use two targeted probes (full code in references/mldrift_shader_probes.md):

  1. In-shader param readback (MLDRIFT_CACHE_PROBE): In ml_drift/delegate/composite/add_values_to_cache_kernel.cc, bypass the X >= update_length gate and overwrite final_value_k with float4(params.Read(0..3)) and final_value_v with float4(params.Read(4..7)). The host [DBG] dump of kv_cache_k_0 head=[...] and kv_cache_v_0 head=[...] then directly prints the exact int32 values the GPU shader read from args.params!
  2. Host-side OperationDef / TensorStorageType dump (MLDRIFT_SLICE_PROBE): In google3/third_party/ml_drift/common/kernels/strided_slice.cc (inside GetStridedSliceCode), print starts.c, ends.c, src.Channels(), and src.GetStorageType() (1 = BUFFER, 3 = TEXTURE_2D) for small int32 tensors (c <= 16).
Classic Root-Cause Signature (b/565413009)

If MLDRIFT_CACHE_PROBE shows params[0..2] are valid (0, 14, 14) while params[3..6] are garbage (±inf in fp16), and [SLICEDBG] shows:

text
starts.c=0 ends.c=3 | src c=7 storage=1 (BUFFER)      -> valid
starts.c=3 ends.c=7 | src c=7 storage=3 (TEXTURE_2D)  -> GARBAGE (unproduced tensor!)

Look at LiteRtOpSelector::ParamTensorToBuffer in ml_drift/delegate/composite/litert_op_selector.cc:

  • Composite ops (add_values_to_cache, runtime_batched_matmul, sdpa_transposed) declare params via AddSrcBuffer (BUFFER storage), whereas standard ops like StridedSlice (rest = param'[3:7]) consume the original tensor via AddSrcTensor (TEXTURE_2D) without consulting replaced_tensors_.
  • Calling model_builder->UpdateOutputTensor(param_tensor, new_param_tensor.id) re-points the producer node's output from param_id to new_param_tensor.id, leaving param_id with no producer in the GpuModel so any non-buffer consumer reads uninitialized GPU memory.
  • Correct fix: Emit model_builder->Copy(param_tensor, new_param_tensor) instead of UpdateOutputTensor. When param_tensor has no other consumers, LinkNodes() in google3/third_party/ml_drift/common/merge_nodes.cc automatically fuses the single-consumer copy back into the producer.

Phase 6: Unit-Test the GpuModel Invariant & Validate End-to-End

  1. Revert all temporary probes (hg revert on kernel/executor files) before running final validation.

  2. Write a hermetic cc_test (see ml_drift/delegate/composite/litert_op_selector_test.cc) that builds a GpuModel via GpuModelBuilder + LiteRtOpSelector and asserts that every in_id of every GpuNode in gpu_model.nodes is present in defined_tensors (graph inputs ∪ const tensors ∪ node outputs):

    bash
    /google/bin/releases/arca9-local-blaze-cli/blaze-for-agents test \
      //ml_drift/delegate/composite:litert_op_selector_test
  3. Validate on device across:

    • Failing model on OpenCL (--backend=gpu)
    • Failing model on WebGPU (--backend=gpu --use_webgpu)
    • A baseline non-sliding-window model (--models=kanana_i4_107_kv_1024 --backend=gpu)

References

Contributions

To contribute or modify this skill or its references, follow the contribution guidelines before making changes.

© google-ai-edge, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy of google-ai-edge/LiteRT.

  • SKILL.md
  • references/contributing.md
  • references/mldrift_shader_probes.md

Open the folder on GitHubat commit 320bfd5

Compare with similar skills

Debugging Litert GPU Accuracy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debugging Litert GPU Accuracy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debugging Litert GPU Accuracy this skillgoogle-ai-edge/LiteRT3.5k—~3.3kAutomated safety check: PassApache-2.0
Smt E2E Dataflow DebuggingGoogleCloudPlatform/DataflowTemplates1.3k—~1.8kAutomated safety check: PassApache-2.0
Debugging MarchatCod-e-Codes/marchat137—~668Automated safety check: NotesMIT
Alefxberg-io/alef100—~1.7kAutomated safety check: PassMIT
Try Fix Alternative Approachdotnet/maui23k—~8.4kAutomated safety check: PassMIT
CanvasBitterbot-AI/bitterbot-desktop2.5k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Smt E2E Dataflow Debugging

    GoogleCloudPlatform/DataflowTemplates

    Debugs logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and comparing source (e.g.

    1.3k GitHub stars~1.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Debugging Marchat

    Cod-e-Codes/marchat

    Diagnoses marchat client and server issues using -doctor, env configuration, and logs.

    137 GitHub stars~668 tokensUpdated 7 days ago
    DevelopmentAuto-check: notes
  • Alef

    xberg-io/alef

    Use Alef correctly for Rust-to-polyglot binding generation. An agent skill from xberg-io/alef.

    100 GitHub stars~1.7k tokensUpdated today
    DevelopmentAuto-check passed
  • Official

    Attempts one alternative fix for a bug, runs the given test command against it and reports what happened, always differing from existing PR fixes.

    23k GitHub stars~8.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Canvas

    Bitterbot-AI/bitterbot-desktop

    Display and control HTML content on connected Bitterbot nodes (Mac, iOS, Android) via the canvas host server.

    2.5k GitHub stars~1.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Dogfood

    callstackincubator/agent-skills

    Official

    Systematically explore and test a mobile app on iOS/Android with agent-device to find bugs, UX issues, and other problems.

    1.7k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed

More from google-ai-edge/LiteRT

  • Litert Ats Op Generator

    google-ai-edge/LiteRT

    Authors, registers, and tests new operator test generators for the LiteRT Accelerator Test Suite (ATS) in litert/test/generators/ and litert/ats/.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check passed

Works with

Questions about Debugging Litert GPU Accuracy

What does Debugging Litert GPU Accuracy do?

Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU). Debugging Litert GPU Accuracy is an agent skill from google-ai-edge/LiteRT. Diagnoses and fixes on-device GPU numerical corruption, garbage generation, zeroed KV caches, and ML Drift delegate lowering bugs for LiteRT and LiteRT-LM models on Android (OpenCL and WebGPU).

When should I use Debugging Litert GPU Accuracy?

Debugging Litert GPU Accuracy fits situations like: A model produces correct output on CPU but generates garbage; wrong outputs on GPU; comparing CPU vs GPU execution op-by-op; state-buffer-by-state-buffer.

How do I install Debugging Litert GPU Accuracy in Claude Code?

Run `npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a claude-code`. Or copy the skill folder (google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy in google-ai-edge/LiteRT) into .claude/skills/debugging-litert-gpu-accuracy in your project. Claude Code loads it when a task matches its description.

How do I install Debugging Litert GPU Accuracy in Codex?

Run `npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a codex`. Or copy the skill folder (google3/third_party/odml/litert/_agents/skills/debugging_litert_gpu_accuracy in google-ai-edge/LiteRT) into .agents/skills/debugging-litert-gpu-accuracy in your project. Codex loads it when a task matches its description.

Can I use Debugging Litert GPU Accuracy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-ai-edge/LiteRT --skill debugging-litert-gpu-accuracy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging-litert-gpu-accuracy, .gemini/skills/debugging-litert-gpu-accuracy, .github/skills/debugging-litert-gpu-accuracy and .opencode/skills/debugging-litert-gpu-accuracy in your project.

What does Debugging Litert GPU Accuracy need to run?

Going by SKILL.md and its folder, Debugging Litert GPU Accuracy needs the command-line tools its instructions call (adb).

Does Debugging Litert GPU Accuracy access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debugging Litert GPU Accuracy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debugging Litert GPU Accuracy use?

Debugging Litert GPU Accuracy is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debugging Litert GPU Accuracy use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Debugging Litert GPU Accuracy?

Skills that share tags, products or a category with Debugging Litert GPU Accuracy: Smt E2E Dataflow Debugging (GoogleCloudPlatform/DataflowTemplates, 1.3k stars), Debugging Marchat (Cod-e-Codes/marchat, 137 stars), Alef (xberg-io/alef, 100 stars) and Try Fix Alternative Approach (dotnet/maui, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debugging Litert GPU Accuracy?

google-ai-edge (a GitHub organization) maintains it in google-ai-edge/LiteRT, which has 3,475 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 10, 2026.

Source: google-ai-edge/LiteRT on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.