Agent skill

Long Audio To Obsidian

by cafe3310 in cafe3310/public-agent-skills

将语音转写项目输出的复杂文件结构整理合并为适合 Obsidian 归档的 Markdown 文档. An agent skill from cafe3310/public-agent-skills.

Apache-2.0Auto-check passedDocuments & Office

Install Long Audio To Obsidian

skills CLI
$ npx skills add cafe3310/public-agent-skills --skill long-audio-to-obsidian -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cafe3310/public-agent-skills long-audio-to-obsidian --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cafe3310/public-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/long-audio-to-obsidian .claude/skills/long-audio-to-obsidian && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
long-audio-to-obsidian
GitHub stars
255
Token cost
~772 tokens
SKILL.md length
165 words
Files
2 (incl. scripts)
Skills in repo
29
Repo updated
First seen
Licence
Apache-2.0

At a glance

将语音转写项目输出的复杂文件结构整理合并为适合 Obsidian 归档的 Markdown 文档. An agent skill from cafe3310/public-agent-skills.

  • Works in 4 steps: 调查与分析 (Survey & Analyze) → 生成合并配置 (Generate Configuration) → 执行合并 (Execute) → …
  • Documents & Office work in your project
  • SKILL.md covers 使用场景, 工作流 (Workflow), 归档标准 (Archiving Standards) and 注意事项
  • Runs Python scripts from its folder

What it does

Long Audio To Obsidian is an agent skill from cafe3310/public-agent-skills. 将语音转写项目输出的复杂文件结构整理合并为适合 Obsidian 归档的 Markdown 文档

Its SKILL.md is about 770 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/doc_merger.py`).

It sits in Documents & Office. It works with Obsidian. The repository describes itself as: personal agent skills for better QoL. The licence is Apache-2.0.

When your agent uses it

  • Documents & Office work in your project

Example prompts

  • “/long-audio-to-obsidian”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 调查与分析 (Survey & Analyze)
  2. 生成合并配置 (Generate Configuration)
  3. 执行合并 (Execute)
  4. 优化命名与生成索引 (Rename & Index)

What it can do on your machine

Read from SKILL.md and the folder at commit 6c45501. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Long Audio To Obsidian loads about 772 tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 165 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~772

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from cafe3310/public-agent-skills at commit 6c45501, republished under its Apache-2.0 licence (© cafe3310). 165 words, ~772 tokens.

Download SKILL.mdSave it as .claude/skills/long-audio-to-obsidian/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
long-audio-to-obsidian
description
将语音转写项目输出的复杂文件结构整理合并为适合 Obsidian 归档的 Markdown 文档
license
Apache-2.0
author
github/cafe3310
depends_on_binary
python3

Long Audio to Obsidian Archiver

此技能将散乱的语音转写项目文件整理成结构清晰的 Markdown 文档,便于在 Obsidian 等知识库中归档。

它采用 "Agent 规划 + 脚本执行" 的模式:

  1. Agent (你):负责理解目录结构、识别文件类型、对分段文件进行逻辑排序(例如按行号 1-200, 201-400),并生成合并计划。
  2. Script:负责根据你的计划,机械地读取文件并生成 Markdown。

使用场景

当用户要求将一个包含多个原始音频、数十个转写分段、以及各种说明文档的目录“整理归档”时使用。

工作流 (Workflow)

1. 调查与分析 (Survey & Analyze)

首先,使用 list_directory (必要时配合 glob) 扫描目标项目目录。 你需要回答以下问题:

  • 哪些是原始录音文本?(通常在 1-原始文件 或类似目录)
  • 哪些是最终输出的分段脚本?(通常在 5-最终输出,文件名包含行号如 _1-200)
    • 关键任务:你需要根据文件名中的数字对它们进行正确排序。
  • 哪些是项目元数据?(日志、术语表、背景要求等)
  • 哪些是衍生知识/问答?(5-最终输出 中非分段的 Markdown 文件)
2. 生成合并配置 (Generate Configuration)

基于你的分析,构建一个 JSON 配置文件。不要直接输出给用户,而是准备写入临时文件。

JSON 结构定义:

json
{
  "output_dir": "目标项目路径/obsidian_archive",
  "tasks": [
    {
      "filename": "1-原始录音汇总.md",
      "title": "原始录音汇总",
      "sections": [
        {
          "title": "文件1.txt",
          "path": "绝对路径或相对路径/文件1.txt",
          "syntax": "text"
        }
      ]
    },
    {
      "filename": "2-最终输出脚本汇总.md",
      "title": "最终输出脚本汇总",
      "sections": [
        {
          "title": "Part 1 (1-200)",
          "path": "路径/output_1-200.txt",
          "syntax": "text"
        },
        {
          "title": "Part 2 (201-400)",
          "path": "路径/output_201-400.txt",
          "syntax": "text"
        }
        // Agent 必须确保这里的顺序是正确的!
      ]
    }
    // ... 继续为“整理要求汇总”和“知识整理汇总”创建任务
  ]
}
3. 执行合并 (Execute)
  1. 使用 write_file 将 JSON 配置保存为临时文件,例如 merge_plan.json。

  2. 调用脚本执行合并:

    bash
    python3 .gemini/skills/long-audio-to-obsidian/scripts/doc_merger.py merge_plan.json

    (注意:请根据实际安装位置调整脚本路径)

  3. 检查脚本输出,确认生成成功。

  4. 删除临时文件 merge_plan.json。

4. 优化命名与生成索引 (Rename & Index)

脚本生成的只是标准文件名的中间产物。为了方便 Obsidian 使用,你需要进行最后一步的人工智能优化:

  1. 确定项目标题:根据原始目录名或核心文件名,提取一个清晰的项目标题(例如 2024-01-01 会议转写 - 项目 A)。
  2. 重命名文件:使用 run_shell_command 的 mv 命令,将 obsidian_archive 中的标准文件重命名为带前缀的格式。
    • 1-原始录音汇总.md -> {标题} 1 原始录音汇总.md
    • 2-最终输出脚本汇总.md -> {标题} 2 最终输出脚本汇总.md
    • ...以此类推。
  3. 生成索引页:
    • 创建一个名为 {标题}.md 的新文件。
    • 内容应包含指向上述所有文件的 Obsidian WikiLinks ([[文件名]])。
    • 如果有关联的外部产品文档(未合并进来的),也可以在此处添加链接。

索引页内容示例:

markdown
相关文档:

[[2024-01-01 会议转写 - 项目 A 1 原始录音汇总]]

[[2024-01-01 会议转写 - 项目 A 2 最终输出脚本汇总]]

...

归档标准 (Archiving Standards)

请按照以下四个维度组织输出文件(对应 JSON 中的 4 个 task):

  1. 原始录音汇总:包含所有原始录音文本。
  2. 最终输出脚本汇总:包含所有经过校对、合并后的分段文本。顺序至关重要。
  3. 整理要求与日志汇总:
    • 工作日志 (0-工作日志.md)
    • 校对术语表 (3-校对和术语表.md)
    • 分段主题 (4-分段主题.md)
    • 所有背景要求文件 (2-要求和信息/*)
  4. 知识整理汇总:
    • 包含所有衍生生成的知识文件(如 问题和回答-xxx.md)。

注意事项

  • 路径处理:JSON 中的 path 推荐使用绝对路径,或者相对于执行 doc_merger.py 时的相对路径。
  • 容错:如果某个预期的目录不存在(例如没有“知识整理”文件),则不要在 JSON 中包含对应的 task,或者创建一个空的 task。
  • 排序:脚本不会自动排序。必须由 Agent 在生成 JSON 时完成所有文件的逻辑排序。

© cafe3310, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/long-audio-to-obsidian of cafe3310/public-agent-skills.

  • SKILL.md
  • scripts/doc_merger.py

Open the folder on GitHubat commit 6c45501

Compare with similar skills

Long Audio To Obsidian next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Long Audio To Obsidian compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Long Audio To Obsidian this skillcafe3310/public-agent-skills255—~772Automated safety check: PassApache-2.0
Obsidian MarkdownAtmosphere/atmosphere3.8k20 repos~1.3kAutomated safety check: PassApache-2.0
Pandoc DOCX TemplateAchuan-2/pandoc_docx_template1.1k—~1kAutomated safety check: PassNone
Mini Wikitrsoliu/mini-wiki116—~3.1kAutomated safety check: PassMIT
Obsidian MarkdownKevRojo/Dulus167—~1.4kAutomated safety check: PassGPL-3.0
MineruNebutra/MinerU-Skill123—~504Automated safety check: PassMIT

Similar skills

  • Obsidian Markdown

    Atmosphere/atmosphere

    Create and edit Obsidian Flavored Markdown with wikilinks, embeds, callouts, properties, and other Obsidian-specific syntax.

    3.8k GitHub starsUsed in 20 repos~1.3k tokens
    Documents & OfficeAuto-check passed
  • Pandoc DOCX Template

    Achuan-2/pandoc_docx_template

    A skill your agent uses when converting Markdown to Word DOCX or DOCX back to Markdown with Pandoc, especially when the output should use the bundled Chinese Word reference templates, heading…

    1.1k GitHub stars~1k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Mini Wiki

    trsoliu/mini-wiki

    Build and maintain a professional, source-traceable project knowledge network.

    116 GitHub stars~3.1k tokensUpdated 6 days ago
    Documents & OfficeAuto-check passed
  • Obsidian Markdown

    KevRojo/Dulus

    Create and edit Obsidian Flavored Markdown with wikilinks, embeds, callouts, properties, and other Obsidian-specific syntax.

    167 GitHub stars~1.4k tokensUpdated 8 days ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    123 GitHub stars~504 tokensUpdated 17 days ago
    Documents & OfficeAuto-check passed
  • Baalda Guide

    naveedharri/baalda

    Answer any question about Baalda (the team second-brain app at baalda.com) in plain, non-technical language — what it is, what it can and cannot do, which file formats it supports (Markdown, images…

    156 GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed

More from cafe3310/public-agent-skills

All 29 skills in this repo
  • Impeccable

    cafe3310/public-agent-skills

    A skill your agent uses when the user wants to design, redesign, shape, critique, audit, polish, clarify, distill, harden, optimize, adapt, animate, colorize, extract, or otherwise improve a…

    255 GitHub stars~4.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Text Watermark Fountain

    cafe3310/public-agent-skills

    A specialized skill for embedding and extracting resilient watermarks in text by manipulating sentence lengths and using Fountain Codes.

    255 GitHub stars~921 tokensUpdated 3 mo ago
    Auto-check passed
  • Obsidian Todo Collector

    cafe3310/public-agent-skills

    从 Obsidian 知识库中扫描指定时间范围内未完成事件,生成/更新未完成事件整理文档. An agent skill from cafe3310/public-agent-skills.

    255 GitHub stars~575 tokensUpdated 3 mo ago
    Auto-check passed
  • Deep Research

    cafe3310/public-agent-skills

    一个全面、自主的深度研究框架。当用户请求对复杂主题、市场调研、技术格局进行深入的多维度调查,或需要大量网页浏览、数据合成和结构化报告的任何任务时,使用此技能。它协调子代理(subagents)并使用基于文件系统的状态管理来防止上下文膨胀。

    255 GitHub stars~996 tokensUpdated 3 mo ago
    Auto-check passed
  • Markdown New

    cafe3310/public-agent-skills

    通过 markdown.new API 将网页、整站或搜索结果转换为干净的 Markdown. An agent skill from cafe3310/public-agent-skills.

    255 GitHub stars~344 tokensUpdated 3 mo ago
    Auto-check passed
  • Oneshot Website

    cafe3310/public-agent-skills

    Generate immersive, one-shot single-file HTML websites with embedded CSS and JS.

    255 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Questions about Long Audio To Obsidian

What does Long Audio To Obsidian do?

将语音转写项目输出的复杂文件结构整理合并为适合 Obsidian 归档的 Markdown 文档. An agent skill from cafe3310/public-agent-skills. Long Audio To Obsidian is an agent skill from cafe3310/public-agent-skills.

When should I use Long Audio To Obsidian?

Long Audio To Obsidian fits situations like: documents & Office work in your project.

How do I install Long Audio To Obsidian in Claude Code?

Run `npx skills add cafe3310/public-agent-skills --skill long-audio-to-obsidian -a claude-code`. Or copy the skill folder (skills/long-audio-to-obsidian in cafe3310/public-agent-skills) into .claude/skills/long-audio-to-obsidian in your project. Claude Code loads it when a task matches its description.

How do I install Long Audio To Obsidian in Codex?

Run `npx skills add cafe3310/public-agent-skills --skill long-audio-to-obsidian -a codex`. Or copy the skill folder (skills/long-audio-to-obsidian in cafe3310/public-agent-skills) into .agents/skills/long-audio-to-obsidian in your project. Codex loads it when a task matches its description.

Can I use Long Audio To Obsidian in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cafe3310/public-agent-skills --skill long-audio-to-obsidian -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/long-audio-to-obsidian, .gemini/skills/long-audio-to-obsidian, .github/skills/long-audio-to-obsidian and .opencode/skills/long-audio-to-obsidian in your project.

What does Long Audio To Obsidian need to run?

Going by SKILL.md and its folder, Long Audio To Obsidian needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Long Audio To Obsidian access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Long Audio To Obsidian safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Long Audio To Obsidian use?

Long Audio To Obsidian is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Long Audio To Obsidian use?

About 772 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Long Audio To Obsidian?

Skills that share tags, products or a category with Long Audio To Obsidian: Obsidian Markdown (Atmosphere/atmosphere, 3.8k stars), Pandoc DOCX Template (Achuan-2/pandoc_docx_template, 1.1k stars), Mini Wiki (trsoliu/mini-wiki, 116 stars) and Obsidian Markdown (KevRojo/Dulus, 167 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Long Audio To Obsidian?

cafe3310 (a GitHub user) maintains it in cafe3310/public-agent-skills, which has 255 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on June 26, 2026.

Source: cafe3310/public-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.