Instrument Data To Allotrope
aws-samples/amazon-bedrock-agents-healthcare-lifesciences
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.
Extract and validate numeric data from chart images, PDF figures, and official source-data files locally and reproducibly, with calibrated axes, compact-scatter recovery plus residual audits…
$ npx skills add franklee16/academic-research-skills --skill thu-digitizer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install franklee16/academic-research-skills thu-digitizer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/franklee16/academic-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/visualization/thu-digitizer-main .claude/skills/thu-digitizer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "thu-digitizer" agent skill from https://github.com/franklee16/academic-research-skills/tree/master/visualization/thu-digitizer-main into .claude/skills/thu-digitizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thu-digitizer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/franklee16/academic-research-skills/tree/master/visualization/thu-digitizer-mainType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add franklee16/academic-research-skills --skill thu-digitizer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install franklee16/academic-research-skills thu-digitizer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/franklee16/academic-research-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/visualization/thu-digitizer-main .agents/skills/thu-digitizer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "thu-digitizer" agent skill from https://github.com/franklee16/academic-research-skills/tree/master/visualization/thu-digitizer-main into .agents/skills/thu-digitizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thu-digitizer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add franklee16/academic-research-skills --skill thu-digitizer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install franklee16/academic-research-skills thu-digitizer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/franklee16/academic-research-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/visualization/thu-digitizer-main .cursor/skills/thu-digitizer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "thu-digitizer" agent skill from https://github.com/franklee16/academic-research-skills/tree/master/visualization/thu-digitizer-main into .cursor/skills/thu-digitizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thu-digitizer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/franklee16/academic-research-skills.git --path visualization/thu-digitizer-main--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add franklee16/academic-research-skills --skill thu-digitizer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install franklee16/academic-research-skills thu-digitizer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/franklee16/academic-research-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/visualization/thu-digitizer-main .gemini/skills/thu-digitizer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "thu-digitizer" agent skill from https://github.com/franklee16/academic-research-skills/tree/master/visualization/thu-digitizer-main into .gemini/skills/thu-digitizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thu-digitizer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install franklee16/academic-research-skills thu-digitizerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add franklee16/academic-research-skills --skill thu-digitizer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/franklee16/academic-research-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/visualization/thu-digitizer-main .github/skills/thu-digitizer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "thu-digitizer" agent skill from https://github.com/franklee16/academic-research-skills/tree/master/visualization/thu-digitizer-main into .github/skills/thu-digitizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thu-digitizer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add franklee16/academic-research-skills --skill thu-digitizer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install franklee16/academic-research-skills thu-digitizer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/franklee16/academic-research-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/visualization/thu-digitizer-main .opencode/skills/thu-digitizer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "thu-digitizer" agent skill from https://github.com/franklee16/academic-research-skills/tree/master/visualization/thu-digitizer-main into .opencode/skills/thu-digitizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thu-digitizer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
thu-digitizerExtract and validate numeric data from chart images, PDF figures, and official source-data files locally and reproducibly, with calibrated axes, compact-scatter recovery plus residual audits…
Thu Digitizer is an agent skill from franklee16/academic-research-skills. Extract and validate numeric data from chart images, PDF figures, and official source-data files locally and reproducibly, with calibrated axes, compact-scatter recovery plus residual audits, grouped/stacked bars, visible-label pie/donut validation, original-pixel aligned lattice composites such as UpSet plots, PDF vector inspection, line tracing, histograms, boxplots, error bars, CSV/JSON evidence, source-data cross-validation, and regression evaluation. Use when a user asks to recover or recreate chart data…
Its SKILL.md is about 8.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 641 other files, including scripts, reference files and assets (for example `.github/workflows/pages.yml`, `README.md` and `agents/openai.yaml`).
It sits in Documents & Office, covering PDF, Accessibility and Machine learning. The repository describes itself as: Comprehensive collection of Claude Code skills for academic research in economics, finance, and social sciences. The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 9a4b2db. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Thu Digitizer loads about 8.5k tokens when it runs, and up to ~30k if it reads all its reference files. Until then it costs about 189 tokens; SKILL.md has 3,938 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from franklee16/academic-research-skills at commit 9a4b2db, republished under its MIT licence (© franklee16). 3,938 words, ~8,535 tokens.
.claude/skills/thu-digitizer/SKILL.md (or your agent's skills folder). This skill also uses 633 other files; get the full folder from GitHub.Recover chart data without silently inventing values. Work locally unless the user explicitly approves a remote OCR or model service.
For every new image or PDF, read references/unified-routing.md and run the unified preflight before choosing a family-specific extractor:
& python scripts/thu_digitizer.py inspect `
--input figure.png --chart-type histogram `
--output-report preflight-report.json `
--output-spec figure-spec.jsonIf the chart type is unknown, omit --chart-type; the router must return needs_chart_type_confirmation and must not authorize numeric extraction. Treat a supplied chart type, full-image panel, appearance estimate, coordinate model, or route as a proposal until its verification field is explicit. Validate a completed spec with scripts/thu_digitizer.py validate-spec before extraction.
The registry in scripts/extractor_registry.py is the single machine-readable map from chart types and input composition to stable, candidate, assisted, unsupported, and refusal routes. Registry presence is not proof of support. Never send an unknown, non-Cartesian, raster-only, or incompatible chart to a convenient XY/PDF extractor merely because it is available.
Keep source identity, original-pixel measurement, calibrated coordinates, non-invention, ambiguity labels, and review evidence as hard requirements. Treat registered scripts as the default deterministic strategy, not as the only strategy that may emit candidate values. Read references/adaptive-execution.md before selecting the execution freedom.
Treat GPT-5.6 Sol and Terra as strong-model profiles when the runtime identifies the current model by either name. A strong profile may produce a separate model_assisted_candidate when a registered implementation is incompatible with the visible grammar, unstable under reasonable verified bounds, contradicted by its overlay, or returns low_confidence/numeric_output_authorized: false. Preserve every deterministic report unchanged; never rewrite a failed script result as successful. For weak or unidentified profiles, retain the registered deterministic or hybrid workflow.
For a weak or unidentified model, or after any missed/extra-mark correction, read references/weak-model-execution.md. Enforce the inventory, locked-exclusion, registered-detector, family-completeness, and original-resolution review loop. Do not let model narration substitute for the coverage ledger or residual audit.
Measure every raster against the exact original file recorded by SHA-256, width, and height. Keep all pixel coordinates in original_raster_pixels. Never measure a chat preview, thumbnail, browser screenshot, resized copy, enlarged crop, overlay, or recreation; use enlargement only for review. Refuse a source-contract mismatch instead of estimating a scale factor. Preserve the original canvas dimensions in every recreation.
For repeated aligned layers such as top bars, row bars, a membership lattice, connectors, and categorical strips, read references/original-pixel-lattice-composites.md and run the registered lattice-composite candidate. Never pass expected row, column, active-node, or intersection counts to detection.
Before designing, benchmarking, or promoting any new extractor, read references/research-quality-baseline.md. Its scope statements, comparative protocol, type-specific metrics, held-out robustness tests, evidence requirements, and promotion gates are mandatory. Do not call a benchmark-only route a supported extractor, and do not claim to exceed WebPlotDigitizer without its documented fair comparison.
When the source is a PDF, read references/vector-pdf-chart-extraction.md before digitizing. Inspect the requested page with scripts/inspect_pdf_vectors.py before rasterizing:
& python scripts/inspect_pdf_vectors.py `
--input paper.pdf --page 8 --output-report vector-inspection.jsonRoute as follows:
vector_paths_detected or mixed_vector_and_raster: visually inspect the target panel, axes, and path geometry. Mixed content is expected; raster strips can coexist with vector points and curves.low_confidence/not_extracted.This is a candidate assisted route, not a stable automatic extractor. It recovers visible plotted geometry only; never infer author raw observations, replicates, error bars, fit parameters, or hidden curve settings. For a log axis, calibrate log10(value) against PDF coordinates rather than applying a linear mapping. Label a newly fitted curve as a refit from extracted points, not the source author's fit.
When an article provides an official XLSX, CSV, archive, or a user supplies ground-truth data, read references/official-source-data-validation.md before comparing values. Source data becomes validation ground truth only after the panel-to-metric, unit, concentration, and summary-statistic mapping is verified.
validated, partial_validated, metric_absent_in_source_workbook, source_unit_unresolved, or not_comparable. Missing source coverage is not an extraction failure.When publishing a gallery case, the primary CSV, static recreation, and interactive recreation must be driven solely by image- or PDF-visible extraction. Do not publish a “source-data mapping” as an extraction case, even when a workbook makes the redraw look more faithful.
not_extracted; never fill it from a source workbook.For concentration-response panels whose markers and paths remain vector objects, use scripts/candidate_digitize_dose_response_pdf.py through extract_dose_response_pdf. Supply a visually verified panel ROI, a verified main-plot ROI, two anchors for the displayed log-concentration coordinate, two y-axis anchors, and separate marker/curve vector colours for every series. Keep a vehicle/control point on a broken-axis segment distinct from the calibrated concentration axis; never assign it an invented molar concentration.
Recover marker centres, only those error-bar endpoints that remain visibly drawn outside the marker, and the authored curve path as separate products. Label the curve curve_path_traced: a PDF path is not evidence for the authors' 4PL parameters. If an official workbook provides mean, SD and N, validate point values against the displayed mean and visible intervals against mean ± SD/sqrt(N) in a separate CSV. Mark intervals hidden entirely by marker glyphs as sem_occluded_by_marker, not extraction failures, and mark missing vehicle rows as vehicle_not_present_in_source_workbook.
The Nature Communications Fig. 4d regression case can be rebuilt with:
& python scripts/build_natcom_dose_response_case.py `
--pdf article.pdf --source-data source-data.xlsx --output-dir evidenceThis route remains candidate until additional marker styles, linear/log axis variants, raster-only inputs, multi-panel layouts, missing-curve/error-bar cases, held-out publications, and matched WebPlotDigitizer comparisons satisfy the binding promotion gates.
scripts/thu_digitizer.py inspect to record input composition and create a FigureSpec template. An unverified route proposal never authorizes values.The stable extractor supports color-distinct histogram bars with calibrated x/y axes. Use scripts/digitize_histogram.py through its extract_histogram API with explicit plot_bounds, two x-axis calibration points, two y-axis calibration points, and the bar color. Its result is calibrated bin edges and heights, not the original raw observations.
For local evidence, run the deterministic companion benchmark:
& python scripts/run_histogram_benchmark.py --output-dir D:\Scratch\thu-digitizer-histogram-benchmarkIt writes CSV, JSON, and overlay evidence for the clean PNG, low-resolution JPEG, and dark PNG fixtures. For a user chart, retain equivalent CSV, JSON, and overlay evidence alongside the calibrated extraction.
For a rectangular heatmap with a readable continuous colour bar, use scripts/candidate_digitize_heatmap.py through extract_heatmap. Supply the verified grid bounds, ordered row/column labels, row and column counts implied by those labels, colour-bar bounds, and its two endpoint values. The route samples each cell away from borders, maps its modal fill colour to the nearest colour-bar row, and separately detects visibly rendered white significance marks.
Treat a cell matching the top or bottom colour-bar colour as interval-censored (clipped_high or clipped_low): the raster cannot distinguish the displayed endpoint from a source magnitude beyond it. Preserve the immutable colour-derived CSV before any official-source comparison. When an official workbook is available, validate it in a separate CSV/report and keep figure/source significance mismatches rather than rewriting the visible extraction.
The Nature Medicine Fig. 4c regression case and its independent SuppData 8 validation can be rebuilt with:
& python scripts/build_natmed_heatmap_case.py `
--input fig4-full.png --source-data supplementary-data.xlsxThis route remains candidate until clean, low-resolution/JPEG, dark/palette-shifted, missing-colour-bar, nonuniform-grid, and held-out publication cases meet the binding promotion gates.
For a line, time-series, ROC/PR, survival, or fitted-curve panel whose visible
stroke is continuous enough to follow, use the continuity candidate in
scripts/digitize_line_chart.py rather than relying on one-pixel-per-sample
colour lookup:
& python scripts/digitize_line_chart.py `
--input figure.png --output-csv extracted.csv --report extraction_report.json `
--trace-mode continuity --plot-bounds 176,63,945,252 `
--x-px-min 176 --x-value-min 0 --x-px-max 945 --x-value-max 850 `
--y-px-min 63 --y-value-min 1.35 --y-px-max 252 --y-value-max 0.4 `
--series curve=#141414 --sample-values 0,25,50,75,100This route fits an affine calibration in transformed value space (including
log10), accepts additional tick anchors, scores anti-aliased colour softly,
and follows the lowest-cost continuous path across candidate runs. It records
coverage, jumps, support, per-sample pixel/value uncertainty, and explicit gap
statuses. A blank or occluded span remains missing; it is never interpolated to
make a complete table. Horizontal borders, gridlines, legends, crossings and
same-colour annotations still require a verified ROI and a series-level review.
The implementation is a candidate until the held-out raster benchmark and a
matched WebPlotDigitizer comparison pass the promotion gates; the legacy
sample mode remains available for backward-compatible runs.
The stable extractor supports colour-distinct, calibrated multi-group boxplots in vertical or horizontal layouts. Use scripts/digitize_boxplot.py through extract_boxplots with explicit plot_bounds, x/y calibration, box_color, line_color, outlier_color, and a verified orientation. It returns only the visible five-statistic summary (q1, median, q3, lower_whisker, upper_whisker) and visible outliers for each group; it does not recover raw samples.
Require distinct fill, line, and outlier colours. Refuse conservatively with low_confidence when a median, whisker cap/spine pairing, box fill, calibration, or a unique automatic orientation is not visibly supported.
For local evidence, run:
& python scripts/run_boxplot_benchmark.py --output-dir D:\Scratch\thu-digitizer-boxplot-benchmarkThe deterministic benchmark renders four groups in vertical and horizontal clean PNG, low-resolution JPEG, and dark PNG variants, plus a missing-median refusal case. It writes calibrated component evidence, CSV, JSON, and overlays before enforcing the quality gates.
For publication boxplots where one series is unfilled, a coloured fill is split by its median, and outlier rings share the structural line colour, use scripts/candidate_digitize_outline_boxplot.py through extract_outline_boxplots. This candidate detects repeated box-width stroke centres, validates ordered Q3/median/Q1 strokes against the two vertical box sides, pairs whisker caps with centre-spine evidence, classifies filled versus unfilled series by interior colour support, and detects hollow outlier rings outside the visible whiskers.
Keep this route separate from the stable fill-component extractor. Require a calibrated linear y-axis, retain every pixel coordinate and stroke-support diagnostic, and generate an overlay before accepting a group. Report a raster-unresolved whisker as coincident with the adjacent quartile only when the report preserves that flag. The output remains the visible five-number summary and visible outliers—not the underlying observations. Promote this route only after held-out publication figures and the binding comparative gates are complete.
The Nature Medicine Fig. 4b regression case can be rebuilt from a retained official image with:
& python scripts/build_natmed_box_case.py --input fig4-full.pngRun the bundled script after axis calibration. Pixel coordinates use the original raster, with (0, 0) at the upper-left.
& python scripts/digitize_line_chart.py `
--input chart.png --output-csv extracted.csv --report extraction_report.json `
--x-px-min 57 --x-value-min 2004 --x-px-max 363 --x-value-max 2014 `
--y-px-min 1 --y-value-min 70 --y-px-max 297 --y-value-max 0 `
--plot-bounds 44,1,375,297 `
--sample-values 2004,2005,2006,2007,2008,2009,2010,2011,2012,2013,2014 `
--series 'biodiversity=#e2070b' `
--series 'ecosystem_services=#bdbdbd' `
--series 'economic_language=#0000dc' `
--color-tolerance 28 --overlay extraction_overlay.pngUse --error-color '#e6e6e6' --error-tolerance 12 --error-min-span 8 only after sampling the actual error-bar color. The report will mark a bar not_extracted if it lacks enough vertical evidence.
For a Cartesian raster scatter panel with compact filled markers, read references/compact-scatter-extraction.md and run scripts/candidate_digitize_scatter.py once per panel. Do not use digitize_line_chart.py, visual counting, or ad hoc connected components for this grammar. Supply verified plot bounds, at least two anchors per axis, and dark, light, or color marker mode. The candidate uses distance-transform peaks to split partially touching markers and excludes thin axes, curves, and text by compact-radius evidence.
& python scripts/candidate_digitize_scatter.py `
--input figure.png --plot-bounds LEFT,TOP,RIGHT,BOTTOM `
--x-anchor XPIXEL1,XVALUE1 --x-anchor XPIXEL2,XVALUE2 `
--y-anchor YPIXEL1,YVALUE1 --y-anchor YPIXEL2,YVALUE2 `
--marker-mode dark `
--output-csv points.csv --report scatter-report.json --overlay scatter-overlay.pngOpen the overlay at original resolution. Accept candidate values only when the report says numeric_output_authorized: true, residual_audit.status is clear, every ring is centred on a visible marker, multi-peak components are plausible touching markers, and suppressed peaks are reviewed. A magenta residual box blocks authorization and is never promoted into the CSV; correct only visibly wrong configuration/exclusions and rerun. Treat script authorization as candidate-level evidence, not final acceptance: if the overlay contradicts the report or reasonable verified bounds produce an unstable point set, preserve that run as failed deterministic evidence and use the adaptive fallback policy. If the panel prints Pearson's R, pass --annotated-pearson-r; use it only as a validation gate, never as a target for adding or removing points. Do not pass an expected point count. Hollow markers, bubbles, dense swarms, and perfectly coincident or fully occluded points are outside this route.
For an UpSet plot or another raster composite with repeated column bars, row bars, and a complete categorical membership grid, create a source-locked configuration and run the registered candidate:
python scripts/candidate_digitize_lattice_composite.py init-config `
--input figure.png --output lattice-config.json
python scripts/candidate_digitize_lattice_composite.py extract `
--input figure.png --config lattice-config.json --output-dir evidenceFill only visibly verified ROIs, one colour or a verified colour list per layer, rendered-size ranges, support thresholds, labels, printed values, and optional independent axis anchors. Prefer actual row/column bars as guides; when bars are separate from the matrix or subpixel, a complete repeated glyph row/column may be declared membership_guides, with the role recorded and inapplicable value-bar geometry validation disabled. The implementation must derive guide centres before consulting semantic array lengths and must classify every row-by-column cell as active, inactive, or ambiguous. Accept numeric output only when numeric_output_authorized: true, the source identity matches, every configured validation passes, and no cell is ambiguous. Never select guide roles or tune thresholds from expected counts. Keep this route candidate; it recovers visible aligned geometry and verified printed values, not hidden records or occluded memberships.
Use scripts/candidate_digitize_bar_chart.py for colour-distinct vertical or horizontal simple/grouped bars and visible stacked segments after verifying the plot bounds, linear value-axis anchors, category centres, layout, baseline, and one colour per series. The output is visible rectangle geometry and calibrated endpoints, not the records summarized by each bar.
When a legend or decorative swatch shares a series colour inside the plot ROI, pass one or more visually verified exclude_regions (CLI: --exclude-region left,top,right,bottom). The report must retain every excluded component and the exclusion rectangles. Never use exclusion regions merely to remove inconvenient candidate bars, and refuse when a legend cannot be separated from plotted geometry.
The Nature Communications Fig. 2b regression case can be rebuilt with scripts/build_natcom_grouped_bar_case.py. Its primary 32-row CSV is raster-derived; retained PDF rectangles are an independent vector validation and do not overwrite raster values. This route remains candidate until additional held-out real-raster cases and the binding WebPlotDigitizer comparison gates are complete.
The bar candidate also bridges one- to two-pixel value-axis gaps caused by
anti-aliasing or gridlines, prefers the uniquely baseline-connected rectangle
when a legend swatch shares a fill colour, and labels a bar whose baseline is
covered by an in-plot overlay as occluded_by_overlay. When a verified
separately coloured error/interval stroke cuts through a pale bar, the route may
use cross-axis-dilated error-colour evidence to reconnect topology across the
stroke. This does not add fill or change the outer measured endpoints; require
verified_occluder_bridging.occluder_role to remain
topology_only_not_numeric_fill.
For a percent stack, retain each visible segment span without normalization.
An expected-total shortfall is separator-consistent only when the measured
internal pixel/value gaps account for it and no gap exceeds the configured
stack-gap tolerance. Keep values_normalized_or_completed: false; a large gap,
top shortfall, missing segment, or unexplained total mismatch remains
low_confidence.
For a pie/donut whose numeric values are explicitly printed and visibly mapped
to distinct sector colours, read references/labelled-pie-donut-extraction.md and run scripts/candidate_digitize_labelled_donut.py. Supply the original-raster contract, panel bounds, group centres, annular bands, named palette, label anchors, and two transcriptions of every declared visible label.
Authorize only matching transcriptions that pass independent annular-geometry validation. Preserve the printed value and printed group sum even when they do not equal 100. Geometry is validation-only and must never supply or normalize a primary label. Unlabeled sectors, angle-only estimates, source observations, exploded/3D pies, gradients, and overlapping sectors remain unsupported.
Use this route for jitter/strip points drawn over pastel bars, especially when a bar outline has the same hue as its points. Read references/scatter-overlay-points.md before extracting this chart family or changing its candidate implementation.
bar_outline_overlap_candidate; do not silently accept it or discard it.visible_marker_candidate for a locally supported centre, and a candidate/ambiguity layer for bar-edge conflicts or merged marker clusters. Record the original-pixel support score, marker colour, crop, and overlay.partial_visible or not_extracted otherwise.Treat a colour template, a local-shape test, and bar geometry as complementary evidence. Re-run the existing scatter benchmark plus a synthetic pastel-bar-with-overlay case before promoting any implementation change; this procedure alone is a candidate workflow, not a stable extractor.
reason_code values and a declared-slot coverage_ledger for bar and visible-label extractors. Use the lattice route's complete Cartesian-cell classification as its equivalent completeness evidence. Do not use an external expected data count to make either structure complete.Read references/report-and-evolution.md and references/research-quality-baseline.md before adding a corrected case or changing extraction logic.
Run scripts/run_synthetic_benchmark.py --output-dir D:\Scratch\thu-digitizer-benchmark before promoting an extraction change. It generates only local, deterministic charts and truth data for clean/low-resolution lines, scatter plots, grouped bars, and stacked bars. Compare the method report rather than choosing one universal technique:
raster_core_continuity candidate (soft colour evidence,
path continuity, and explicit gaps).scripts/run_scatter_benchmark.py.Read references/benchmark-suite.md before interpreting or expanding this suite.
The grouped/stacked-bar routes in this synthetic runner are benchmarks, not dedicated stable user-facing extractors. Do not claim automated bar support until a dedicated extractor meets the binding research-quality baseline.
Use the R Graph Gallery as a public taxonomy and rendering-pattern reference, not as a source of truth data. Read references/r-graph-gallery-benchmark.md before adding its representative families. Generate fresh local data and renderings; retain only page URLs and family metadata unless the user separately approves keeping downloaded site assets or source code.
Begin with exact-value families (line/scatter, bars, histogram, area, heatmap, box plot). For density/violin, unlabeled pie/donut, general polar, treemap, maps, networks, Sankey, word clouds, and animation, state the recoverable representation explicitly; a raster cannot generally recover the original raw observations or hidden graph structure. Use the assisted labelled pie/donut route only for explicitly printed values with independent visible-sector validation.
Keep the stable skill code unchanged during normal extraction. Record a candidate improvement only when a correction identifies a reproducible failure mode. Compare the candidate against the current implementation on all benchmark cases; promote it only after it improves the relevant metric without regressions and after user approval. Keep sensitive source images outside the skill by default.
The bundled executables support calibrated colour-distinct lines plus pale error bars, a continuity-aware raster line candidate, a candidate compact filled-scatter route with distance-peak splitting and a blocking residual audit, a candidate original-pixel aligned-lattice route, grouped/stacked bar candidates with verified-occluder topology bridging, an assisted visible-label pie/donut candidate, calibrated histograms, colour-distinct vertical or horizontal boxplots, and PDF vector inspection. The compact-scatter route is not a universal point detector: hollow markers, bubbles, dense swarms, perfect coincidences, and occluded marks remain unsupported or not_extracted. The pie/donut route does not discover OCR labels or infer unlabeled sectors. The lattice route does not cover irregular matrices, merged or occluded cells, or unverified semantic text. Direct PDF marker recovery and source-data validation are assisted workflows. Histogram output is limited to visible bin edges/heights and boxplot output to visible summaries/outliers. For bars, OCR-heavy images, nonlinear or unsupported coordinates, overlapping series, or unverified vector paths, retain the same evidence/refusal gates and do not claim stable automation without dedicated held-out evidence.
© franklee16, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 633 other files (scripts, references, assets) in visualization/thu-digitizer-main of franklee16/academic-research-skills.
Open the folder on GitHubat commit 9a4b2db
Thu Digitizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Thu Digitizer this skillfranklee16/academic-research-skills | 223 | — | ~8.5k | Automated safety check: Pass | MIT | |
| Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences | 274 | 2 repos | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Uap Release Analyzerckpxgfnksd-max/uap-release-analyzer | 155 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Research Integrity Auditxuzhougeng/wisp-science | 1k | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 | |
| File ReadingWide-Moat/open-computer-use | 126 | 1 repos | ~3.1k | Automated safety check: Pass | Proprietary | |
| CSV To Executive Reportskrun-dev/skrun | 210 | — | ~1.1k | Automated safety check: Pass | MIT |
aws-samples/amazon-bedrock-agents-healthcare-lifesciences
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.
ckpxgfnksd-max/uap-release-analyzer
Inventory, extract, and analyze tranches of declassified UAP/UFO files — including war.gov/UFO/ "PURSUE" releases, FBI Vault, NARA boxes, and AARO publications.
xuzhougeng/wisp-science
学术审查 / research-integrity screening of a manuscript's figures and reported numbers.
Wide-Moat/open-computer-use
A skill your agent uses when a file has been uploaded but its content is NOT in your context — only its path at /mnt/user-data/uploads/ is listed in an uploadedfiles block.
skrun-dev/skrun
Turn a CSV of operational data (sales, usage, signups, support tickets) into a multi-page styled PDF executive report with narrative + matplotlib charts.
ComPDFKit/compdf-skills
Convert Word, Excel, PPT, HTML, TXT, CSV, RTF, PNG, and JPG files into PDF with ComPDF.
franklee16/academic-research-skills
End-to-end econometric analysis and economics/management paper-writing workflow.
franklee16/academic-research-skills
Comprehensive workflow for handling journal Revise and Resubmit (R&R) decisions.
franklee16/academic-research-skills
A skill your agent uses when researchers need Chinese academic prose translated into publication-oriented English or English manuscript paragraphs and complete sections polished for SCI, SSCI, or…
franklee16/academic-research-skills
Assess and monitor an ongoing research project's competitor landscape, novelty risk, and method/data opportunities from a proposal, paper, research question, data or method notes, or a suspected…
franklee16/academic-research-skills
Generate or fill in academic grant application forms (project statement, education plan, pathway to impact, references) using draft research material.
franklee16/academic-research-skills
A skill your agent uses when a user needs通用中文参考文献、GB/T 7714-style bibliography entries, Chinese academic reference formatting, or BibTeX completion from Chinese or English literature titles…
Categories
Extract and validate numeric data from chart images, PDF figures, and official source-data files locally and reproducibly, with calibrated axes, compact-scatter recovery plus residual audits…. Thu Digitizer is an agent skill from franklee16/academic-research-skills. Extract and validate numeric data from chart images, PDF figures, and official source-data files locally and reproducibly, with calibrated axes, compact-scatter recovery plus residual audits, grouped/stacked bars, visible-label pie/donut validation, original-pixel aligned lattice composites such as UpSet plots, PDF vector inspection, line tracing, histograms, boxplots, error bars, CSV/JSON evidence, source-data cross-validation, and regression evaluation.
Thu Digitizer fits situations like: A user asks to recover; recreate chart data; extract scatter points; printed pie/donut labels.
Run `npx skills add franklee16/academic-research-skills --skill thu-digitizer -a claude-code`. Or copy the skill folder (visualization/thu-digitizer-main in franklee16/academic-research-skills) into .claude/skills/thu-digitizer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add franklee16/academic-research-skills --skill thu-digitizer -a codex`. Or copy the skill folder (visualization/thu-digitizer-main in franklee16/academic-research-skills) into .agents/skills/thu-digitizer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add franklee16/academic-research-skills --skill thu-digitizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/thu-digitizer, .gemini/skills/thu-digitizer, .github/skills/thu-digitizer and .opencode/skills/thu-digitizer in your project.
Going by SKILL.md and its folder, Thu Digitizer needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Thu Digitizer is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.5k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Thu Digitizer: Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Uap Release Analyzer (ckpxgfnksd-max/uap-release-analyzer, 155 stars), Research Integrity Audit (xuzhougeng/wisp-science, 1k stars) and File Reading (Wide-Moat/open-computer-use, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
franklee16 (a GitHub user) maintains it in franklee16/academic-research-skills, which has 223 GitHub stars. The repository holds 1,617 skills in this directory. The repository was last updated on September 18, 2026.
Source: franklee16/academic-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.