Apify Trend Analysis
sickn33/agentic-awesome-skills
Discover and track emerging trends across Google Trends, Instagram, Facebook, YouTube, and TikTok to inform content strategy.
Search the user's local Pensieve screenshot archive by text, app, or time range.
$ npx skills add arkohut/pensieve --skill pensieve-search -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install arkohut/pensieve pensieve-search --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/arkohut/pensieve.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pensieve-search .claude/skills/pensieve-search && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pensieve-search" agent skill from https://github.com/arkohut/pensieve/tree/master/skills/pensieve-search into .claude/skills/pensieve-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pensieve-search", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/arkohut/pensieve/tree/master/skills/pensieve-searchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add arkohut/pensieve --skill pensieve-search -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install arkohut/pensieve pensieve-search --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/arkohut/pensieve.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/pensieve-search .agents/skills/pensieve-search && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pensieve-search" agent skill from https://github.com/arkohut/pensieve/tree/master/skills/pensieve-search into .agents/skills/pensieve-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pensieve-search", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add arkohut/pensieve --skill pensieve-search -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install arkohut/pensieve pensieve-search --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/arkohut/pensieve.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/pensieve-search .cursor/skills/pensieve-search && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pensieve-search" agent skill from https://github.com/arkohut/pensieve/tree/master/skills/pensieve-search into .cursor/skills/pensieve-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pensieve-search", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/arkohut/pensieve.git --path skills/pensieve-search--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add arkohut/pensieve --skill pensieve-search -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install arkohut/pensieve pensieve-search --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/arkohut/pensieve.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/pensieve-search .gemini/skills/pensieve-search && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pensieve-search" agent skill from https://github.com/arkohut/pensieve/tree/master/skills/pensieve-search into .gemini/skills/pensieve-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pensieve-search", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install arkohut/pensieve pensieve-searchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add arkohut/pensieve --skill pensieve-search -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/arkohut/pensieve.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/pensieve-search .github/skills/pensieve-search && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pensieve-search" agent skill from https://github.com/arkohut/pensieve/tree/master/skills/pensieve-search into .github/skills/pensieve-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pensieve-search", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add arkohut/pensieve --skill pensieve-search -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install arkohut/pensieve pensieve-search --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/arkohut/pensieve.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/pensieve-search .opencode/skills/pensieve-search && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pensieve-search" agent skill from https://github.com/arkohut/pensieve/tree/master/skills/pensieve-search into .opencode/skills/pensieve-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pensieve-search", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pensieve-searchSearch the user's local Pensieve screenshot archive by text, app, or time range.
Pensieve Search is an agent skill from arkohut/pensieve. Search the user's local Pensieve screenshot archive by text, app, or time range. Use when the user asks to find a screenshot ("find that thing I looked at last week", "show me when I was working on the Mastra integration", "find screenshots of the YouTube video about X"), or to locate a specific moment in time across captured activity. Returns ranked entityids with filepaths, timestamps, OCR snippets, and VLM-extracted structured metadata (app, topic, workspace). Can open the image locally for the user. Skip when…
Its SKILL.md is about 8.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Forecasting and time series. It works with Mastra and YouTube. The repository describes itself as: A passive recording project allows you to have complete control over your data. Automatically take screenshots of all your screens, index them, and save them locally. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3e55c66. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqcurlpython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
bilibili.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pensieve Search loads about 8.2k tokens when it runs. Until then it costs about 162 tokens; SKILL.md has 3,238 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from arkohut/pensieve at commit 3e55c66, republished under its Apache-2.0 licence (© arkohut). 3,238 words, ~8,191 tokens.
.claude/skills/pensieve-search/SKILL.md (or your agent's skills folder).Pensieve (memos) continuously captures screenshots and indexes them. This skill teaches you to query that index over its HTTP API.
http://127.0.0.1:8839 by default (config: server_host / server_port).MEMOS_SERVER_HOST / MEMOS_SERVER_PORT env vars or ~/.memos/config.yaml if the user has a non-default setup. Confirm with the user if the default doesn't respond.~/.memos/screenshots/YYYYMMDD/<entity>.webp. The API returns filepath directly so you don't have to compute it.http://127.0.0.1:8839/?q=...&start=...&end=...&app_names=...&library_ids=... — you should always offer this URL in your reply so the user can click and verify results visually.http://127.0.0.1:8839/entities/<id> — opens one specific screenshot in context with its metadata. Offer this for any hit you cite.http://127.0.0.1:8839/config — surface this if the user asks about settings.Before the first query in a session, verify the server is up:
curl -s -o /dev/null -w "%{http_code}\n" "http://127.0.0.1:8839/api/search?q=test"200 means good. 000 or connection-refused means memos isn't running — tell the user to start it (memos serve or system-service).
Always quote the URL — the ? and & are zsh glob characters and will fail with no matches found if unquoted. Every curl example below assumes a quoted URL.
GET /api/searchParameters:
| Param | Type | Notes |
|---|---|---|
q | string | Full-text query (jieba-tokenized; works for Chinese + English mixed). Empty q="" returns recent screenshots without ranking. |
start, end | int (UNIX seconds, UTC) | Time window on file_created_at. Both required together. |
app_names | string | Comma-separated active_app values, e.g. "Google Chrome,iTerm2". |
library_ids | string | Comma-separated; usually leave unset to query all. |
limit | int 1..200 | Default 48. Use 10–20 for casual queries, 200 only when scanning. |
facet | bool | Include facet_counts (per-app counts), date_range (earliest/latest), and date_buckets (count per day or month, see Strategy 5). Set true on broad queries to help narrow. |
date | string | Bucket filter, YYYY-MM or YYYY-MM-DD. Intersected with start/end. |
Returns SearchResult { hits: [...] }. Each hit:
{
"document": {
"id": "12345",
"filepath": "/Users/.../screenshots/20260423/00001.webp",
"file_created_at": "2026-04-23T04:18:21+00:00",
"tags": [],
"metadata_entries": [
{"key": "active_app", "value": "iTerm2", "source": "..."},
{"key": "active_window", "value": "✳ Debug memos service background job issue", "source": "..."},
{"key": "ocr_result", "value": "...long JSON of OCR boxes...", "source": "ocr"},
{"key": "structured_vlm_v1_qwen3_6_35b", "value": {"primary": {"app": "Claude Code", "title_or_topic": "...", "what": "...", "workspace": "memos"}}, "source": "structured_vlm"},
...
]
}
}The metadata_entries is the goldmine — the VLM-extracted structured fields (primary.app, primary.title_or_topic, primary.what, primary.workspace) live inside the structured_vlm_v1_* entry's value. In /api/search hits the value is a parsed JSON object, not a string — access fields directly with .value.primary.what, and do NOT pipe through fromjson (it errors with "only strings can be parsed"). Note the inconsistency: /api/entities/{id} and the /context endpoint (Strategy 6) return the same value as a JSON string, where you do need fromjson. The portable branch is (.value | if type=="string" then fromjson else . end). The exact key suffix (e.g. qwen3_6_35b) varies per install — match with select(.key | startswith("structured_vlm")).
Important: that JSON metadata is also fully tokenized into the FTS index. So q="Claude Code Mastra workspace" will match a screenshot whose VLM said app=Claude Code, what="...Mastra...", workspace=... — even if "Mastra" doesn't appear in the OCR text.
Resist the urge to pre-anchor on what you think the user means. A "find 最终幻想" question sounds like it's about the game, but the same brand lives across YouTube videos, Bilibili clips, wallpaper engines, wikis, store pages, and game launchers. The user's casual phrasing names the topic, not the surface — only the corpus knows where they actually consumed it.
So run one wide query first and let the results tell you what surfaces exist:
curl -s --get 'http://127.0.0.1:8839/api/search' \
--data-urlencode 'q=最终幻想' --data-urlencode 'facet=true' --data-urlencode 'limit=20' > /tmp/raw.json
jq -r '.hits[].document
| "\(.file_created_at) \([.metadata_entries[] | select(.key=="active_app") | .value][0] // "?") \([.metadata_entries[] | select(.key=="url") | .value // "-"][0]) \([.metadata_entries[] | select(.key=="active_window") | .value // "-"][0] | .[0:80])"' /tmp/raw.jsonSample output (verified on a real install):
2026-04-03T15:26Z Google Chrome about:blank (55) When Tifa played「Tifa's T…Rebirth🖤 Ru's Piano - YouTube
2026-04-02T15:47Z Google Chrome about:blank (55) Final Fantasy X「To Zanark… Medley | Ru's Piano - YouTube
2026-02-15T15:59Z Google Chrome https://www.bilibili.com/video/BV1K2HQzbEAy… “蒂法的审视 手机动态壁纸 wallpaper engine_哔哩哔哩_bilibili”🔊
2026-01-30T05:17Z Google Chrome https://www.bilibili.com/ 哔哩哔哩 (゜-゜)つロ 干杯~-bilibiliNow you know: the user doesn't play Final Fantasy — they watch Ru's Piano play FF themes on YouTube, and 蒂法的审视 wallpaper engine clips on Bilibili. If you'd anchored on q=最终幻想 game or guessed app_names=Final Fantasy you'd have found nothing. The corpus revealed both the surfaces (YouTube + Bilibili) AND the user's actual relationship to the topic (music + wallpapers, not gameplay).
Now you can build per-surface queries anchored on each surface's stable identifier and merge the timelines.
Anchor on stable identifiers, not free-text names, because free-text is noisy — short or common substrings match the library_ids= parameter inside memos's own URL bar, and a CJK brand name often matches OCR'd filenames of unrelated downloaded videos. Stable identifiers come from the URL or the exact app name:
guancha.cn, youtube.com/watch?v=…)space.bilibili.com/<uid> (e.g. 10330740)@mastra-ai, @RusPiano)active_app value (iTerm2, Google Chrome, 企业微信)Cross-check every match by filtering hits client-side on the URL or app fields — never trust the FTS hit list alone for "did the user actually visit / consume X" questions. The skill's job is to use multiple raw queries + client-side filtering to give the user a true answer, not to take the first FTS rank as gospel.
curl -s 'http://127.0.0.1:8839/api/search?q=Mastra+roadmap&limit=15' | jq '.hits[].document | {id, filepath, file_created_at}'Joining terms with + (URL-encoded space) gives an AND query (jieba splits, then FTS5 does AND-of-tokens). Mix English + Chinese freely: q=理赔+Demo+演示.
The user's active_app is recorded literally per OS. Common values:
Google Chrome, iTerm2, Claude, Cursor, WeChat, 企业微信msedge.exe, chrome.exe, WeChat.exe, claude.execurl -s 'http://127.0.0.1:8839/api/search?q=mastra+roadmap&app_names=Google+Chrome&limit=10' \
| jq '.hits[].document | {id, filepath, file_created_at}'Times are UNIX seconds, UTC. Compute with date:
# "last 7 days"
SINCE=$(date -v-7d -u +%s 2>/dev/null || date -d '7 days ago' -u +%s)
NOW=$(date -u +%s)
curl -s "http://127.0.0.1:8839/api/search?q=mastra&start=$SINCE&end=$NOW&limit=20"For specific local dates: parse the user's date words (e.g. "上周三" → resolve to absolute → 00:00 local → UNIX UTC) and use a 1- or 24-hour window.
The fastest path to a precise hit is q + app + tight time window. For "find the YouTube Mastra roadmap livestream from last week":
SINCE=$(date -v-10d -u +%s 2>/dev/null || date -d '10 days ago' -u +%s)
NOW=$(date -u +%s)
curl -s "http://127.0.0.1:8839/api/search?q=mastra+roadmap+youtube&app_names=Google+Chrome&start=$SINCE&end=$NOW&limit=10" > /tmp/hits.json
jq '.hits[].document | {id, filepath, file_created_at}' /tmp/hits.jsonTip: pipe the response to a tmp file before complex jq filters. Inline shell-quoted jq with nested \"...\" escapes is fragile — use a file.
When q returns hundreds of results, set facet=true to see what apps and time range dominate the matches:
curl -s 'http://127.0.0.1:8839/api/search?q=mastra&facet=true&limit=5' \
| jq '{date_range, bucket_unit, date_buckets: .date_buckets[:10], facet_counts: .facet_counts[0].counts[:10]}'date_range (top-level): {earliest, latest} ISO timestamps spanning all matched entities under the current filters. Use it to suggest a tighter start/end window.date_buckets (top-level): [{date, count}] grouped by day or month — the bucket unit is adaptive based on the matched span (≤ 60 days → day, else month). The chosen unit is reported in bucket_unit. Single-bucket results are suppressed (returned as empty + bucket_unit=null) since they're not useful for narrowing.facet_counts[0].counts: list of {value, count} for active_app, sorted desc. Use it to suggest an app_names= filter, or to ask the user which app they meant.All three are populated only when facet=true (or when settings.facet=true server-side).
To drill into a bucket, re-issue with ?date=YYYY-MM or ?date=YYYY-MM-DD. The server intersects date with the existing filters; if you also set start/end, the effective range is the overlap. After drilling into a month, the next response's date_buckets automatically adapts to days within that month.
grep -C for the timeline)Search pinpoints a moment; it doesn't show what surrounds it. When a hit lands on the right frame but the answer isn't on that frame — the working directory a command ran in, what led up to it, what happened next — pull its temporal neighbors. Think grep -C N for the screenshot timeline, especially the before side where the lead-up lives (the cwd, the command typed, the app you came from). This is what the web UI's context strip does.
GET /api/libraries/{library_id}/entities/{entity_id}/context?prev=N&next=Nprev / next are counts (0–100 each), ordered by file_created_at — true temporal neighbors. Robust where id arithmetic is not: entity ids are sparse and non-contiguous, so "fetch ids around X" silently misses frames. Always use this endpoint, never id math.{ "prev": [entity…], "next": [entity…] }. The target is not included — you already have it (the id you passed), sitting chronologically between the last prev and the first next.library_id comes from the entity itself: curl …/api/entities/{id} | jq -r .library_id (or reuse LIB). Frontend default (web/src/lib/api/entities.ts → fetchEntityContext) is a symmetric ?prev=12&next=12 (grep -C 12) — match it.Three gotchas, all differing from /api/search:
ocr_result. A prev=12&next=12 call is ~230 KB; prev=100&next=100 is multi-MB. Project only the fields you need (jq below) — never dump the raw response.2026-06-07T01:51:51, no Z/offset. They are UTC; append Z (or convert) before you cite a time, or you'll report the wrong hour.structured_vlm value is a JSON string, not a parsed object — fromjson it to reach .primary.what.Recipe. Get library_id, then project just the fields the question needs — that projection is the jq's only job (it keeps the bloated response small):
B=http://127.0.0.1:8839; id=1748373
lib=$(curl -s "$B/api/entities/$id" | jq -r .library_id)
curl -s "$B/api/libraries/$lib/entities/$id/context?prev=12&next=12" | jq -r '
.prev[], .next[] | "\(.id)\t\(.file_created_at)Z\t"
+ ([.metadata_entries[]|select(.key=="active_window").value][0] // "-") + "\t"
+ ([.metadata_entries[]|select(.key|startswith("structured_vlm")).value][0] // "{}" | fromjson.primary.what // "-")'Don't capture the body with $(curl …) — OCR/VLM text mangles JSON in command substitution; pipe straight to jq (or a tmp file).
Worked example — the directory a codex command ran in. Search pinpointed the codex launch (entity 1748373), but that frame was full-screen with the version-update prompt — no directory visible. Pull its neighbors (the recipe above, id=1748373):
1748372 2026-06-07T01:51:45Z you@host:~/projects/memos/pensieve-rpa 在终端中执行 'z pensieve' 命令切换目录并准备运行…
1748374 2026-06-07T01:51:56Z ✳ Review blocking feature project for… …The frame 6 s before the launch (1748372) shows the shell prompt …:~/projects/memos/pensieve-rpa and a VLM note that the user ran z pensieve to jump directories — recovering the answer (~/projects/memos/pensieve-rpa) that wasn't on the target frame at all. (Target 1748373 itself isn't in the strip; you supply it.)
Widening (mirrors clicking an edge thumbnail in the UI). If the lead-up frame isn't in the window yet: bump the counts (prev=30&next=30 → 60 → 100, the per-side max), or re-center on the oldest frame returned and pull its context — the same walk-further-back the UI does, keeping each response small so you can scan back indefinitely.
Citing it in a reply. When context resolves the answer, cite the lead-up frame's entity URL next to the pinpoint hit, e.g. "ran in ~/projects/memos/pensieve-rpa — the prompt is visible 6 s earlier → http://127.0.0.1:8839/entities/1748372". The user clicks through to verify the exact frame.
/api/search caps limit at 200 per request, but the response now exposes the real total under filters in the found field. Use it directly — no heuristics needed.
found and out_ofEvery response includes:
{ "found": 2883, "out_of": 1464297, "hits": [ ... 200 items ... ], ... }found — unbounded count of FTS matches under your filters (q, start/end, app_names, library_ids). This is the truthful answer to "how many things matched my keywords?". If found > limit, there are more matches you didn't get back.out_of — total entities in the collection scope (library_ids only — q/start/end/app_names are dropped). Typesense convention: "found N out of M total". Surface this only when the user asks dataset-level questions ("how many screenshots do I have?"); for normal "find X" replies use found.Edge case: very rarely, found < len(hits) when vector search contributes hits whose FTS rank was zero (e.g. queries with no good keyword match but semantic neighbors). Treat found as the keyword-match count, with the understanding that hits may include a few extra semantic neighbors.
resp = search(q, start, end, limit=200)
hits = resp.hits
found = resp.found
if found <= len(hits):
return hits # complete, done
if user wants "top match" / "first few":
return hits[:N] # top-ranked are already here
if user wants comprehensive scan ("show me everything ..."):
return time_slice(q, start, end) # recipe below — still needed (no offset yet)
else:
tell user: f"{found} matches, narrow with app or shorter time window"For "show me every screenshot of X in the last week", recursively halve the time window when truncation hits. Always use a tmp file for the response body — bash command substitution $(curl ...) mangles JSON that contains literal newlines or control chars (which OCR / VLM metadata frequently does).
# bash function — paste into agent shell or save as ~/bin/pensieve-paged-search.sh
paged_search() {
local q="$1" start="$2" end="$3"
local tmp
tmp=$(mktemp)
curl -s "http://127.0.0.1:8839/api/search?q=$(echo -n "$q" | jq -sRr @uri)&start=$start&end=$end&limit=200" > "$tmp"
local n
n=$(jq '.hits | length' "$tmp")
if [ "$n" -lt 200 ] || [ $((end - start)) -lt 60 ]; then
jq -c '.hits[]' "$tmp"
rm "$tmp"
return
fi
rm "$tmp"
local mid=$(( (start + end) / 2 ))
paged_search "$q" "$start" "$mid"
paged_search "$q" "$mid" "$end"
}
# Usage: collect all 'memos' hits in the last 7 days
SINCE=$(date -v-7d -u +%s 2>/dev/null || date -d '7 days ago' -u +%s)
NOW=$(date -u +%s)
paged_search "memos" "$SINCE" "$NOW" > /tmp/all_hits.ndjson
wc -l /tmp/all_hits.ndjson # raw hit count (may include boundary dups)Boundary duplicates: the mid second may match in both halves if multiple captures share that timestamp. Dedupe by id:
jq -s 'unique_by(.document.id)' /tmp/all_hits.ndjson > /tmp/dedup.json
jq 'length' /tmp/dedup.jsonIn testing, a "last 24h" query for q=memos produced 2888 raw hits → 2883 unique. Boundary dup rate is small but real.
If a 60-second window still saturates 200 hits, the burst is too dense to enumerate (e.g. continuous scrolling capturing every 4 s = 15 captures/min × 5+ apps multiplied = saturation). Stop recursion there and report: "this minute hit the cap; the screen activity was too dense to enumerate, please narrow further".
Order matters — hybrid_search ranks by reciprocal rank fusion of FTS (weight 0.7) + vector embedding similarity (0.3). The top hit is most likely the right one for direct lookups; for "find all" queries, walk all hits.
Per hit, extract the human-readable summary. Save to a tmp file first to keep the jq filter readable:
curl -s 'http://127.0.0.1:8839/api/search?q=mastra&limit=10' > /tmp/hits.json
# Per-hit one-liner: timestamp + app + topic + truncated 'what'
jq -r '.hits[].document |
"\(.file_created_at) \([.metadata_entries[]
| select(.key | startswith("structured_vlm"))
| .value.primary
| "app=\(.app // "?") topic=\(.title_or_topic // "?") what=\((.what // "?")[0:80])"][0])"' /tmp/hits.jsonSample output (verified):
2026-04-30T01:35:21Z app=Google Chrome topic=mastra roadmap what=在 Google 搜索框中输入并搜索 'mastra roadmap'
2026-04-30T01:35:26Z app=Google Chrome topic=mastra roadmap what=在 Google 搜索框中输入并搜索 'mastra roadmap'
2026-05-03T13:03:00Z app=iTerm2 topic=Agent 检索 what=在终端内运行 Claude Code 进行代码开发...primary.what is in the user's interface language (Chinese for Chinese-locale machines). Don't translate unless asked — show it as-is.
Timestamp is file_created_at (ISO 8601 UTC). Convert to user-local for display.
After every search, return both:
Mirror the API params, URL-encoded. Always include both q and submitted_q with the same value — q populates the search input, submitted_q activates the facet sidebar with the right counts. Sharing only q leaves the facets blank until the user hits Enter.
app_names and library_ids use repeated-key style in the web URL (&app_names=A&app_names=B); but the route's z.array(z.coerce.number()) schema also accepts comma-separated.
# Build a search-page URL for a query
QUERY="mastra roadmap"
SINCE=1746230400
NOW=1746834400
APPS="Google Chrome"
python3 -c "
import urllib.parse
q = urllib.parse.quote('$QUERY')
app = urllib.parse.quote('$APPS')
print(f'http://127.0.0.1:8839/?q={q}&submitted_q={q}&start=$SINCE&end=$NOW&app_names={app}')
"
# → http://127.0.0.1:8839/?q=mastra%20roadmap&submitted_q=mastra%20roadmap&start=1746230400&end=1746834400&app_names=Google%20ChromeENTITY_ID=1646223
echo "http://127.0.0.1:8839/entities/$ENTITY_ID"When the user asks "find X", reply in roughly this shape:
Found **3 of 47 matches** for **mastra roadmap** in the last 7 days:
1. **2026-04-30 09:35 (Google Chrome)** — Searching "mastra roadmap" in Google
→ http://127.0.0.1:8839/entities/1634834
2. **2026-04-30 09:35 (Google Chrome)** — Same search, second frame
→ http://127.0.0.1:8839/entities/1634835
3. **2026-05-03 21:03 (iTerm2)** — Claude Code session discussing agent retrieval for the Mastra livestream
→ http://127.0.0.1:8839/entities/1643716
[See all in browser](http://127.0.0.1:8839/?q=mastra%20roadmap&submitted_q=mastra%20roadmap&start=1746230400&end=1746834400)Always include the [See all in browser](...) link, even when there's only 1 hit — it lets the user re-run the query and tweak filters without going through you. Always include the per-hit entities/<id> URL because clicking it opens the full-resolution screenshot with surrounding metadata, far richer than what you can fit in a CLI reply.
Show found when it exceeds returned hits (e.g. "3 of 47 matches"), so the user knows there's more to browse. If found == len(hits), just say "Found 3 hits ..." without the total. out_of is collection size (matches Typesense semantics) — surface it only when the user asks "how many screenshots do I have total?" or similar dataset-level questions.
For paginated/sliced results (the paged_search recipe earlier), still link to the unsliced search page URL — the user wants to browse, not to see your slicing internals.
Why this matters: search quality depends on the structured_vlm_v1_* metadata field on each entity. That field is what holds primary.app / primary.what / primary.workspace — the LLM-readable summary that makes "find Mastra roadmap" work even when "Mastra" isn't visible in the OCR text. Entities lacking this field only have OCR; their search relevance is much lower.
Two sources of gaps:
The plugin pipeline is idempotent and re-runnable, so backfilling is safe.
curl -s http://127.0.0.1:8839/api/plugins | jq '.[] | select(.name == "builtin_structured_vlm") | .id'
# → 3 (varies per install)A user can have multiple libraries (test imports, archives). The "live" one is named per default_library in config:
DEFAULT_LIB_NAME=$(curl -s http://127.0.0.1:8839/api/config | jq -r '.default_library')
curl -s http://127.0.0.1:8839/api/libraries \
| jq --arg n "$DEFAULT_LIB_NAME" '.[] | select(.name == $n) | {id, folders: [.folders[] | {id, path}]}'
# → {"id": 6, "folders": [{"id": 14, "path": "/Users/.../.memos/screenshots"}]}Capture LIB and FOLDER ids and the screenshots_root path for use below.
Walk a day's entities, count those missing structured_vlm in plugin_status. Use a tmp dir for the JSON bodies (same control-char gotcha as in pagination).
LIB=6; FOLDER=14; PLUGIN=3
SCREENSHOTS_ROOT="$HOME/.memos/screenshots"
DAY=20260401
TMPDIR=$(mktemp -d)
audit_day() {
local day="$1"
local has=0 miss=0 offset=0
while :; do
curl -s "http://127.0.0.1:8839/api/libraries/$LIB/folders/$FOLDER/entities?limit=400&offset=$offset&path_prefix=$SCREENSHOTS_ROOT/$day" > "$TMPDIR/batch.json"
local n
n=$(jq 'length' "$TMPDIR/batch.json")
[ "$n" -eq 0 ] && break
local h
h=$(jq "[.[] | select(.plugin_status | map(.plugin_id) | index($PLUGIN))] | length" "$TMPDIR/batch.json")
has=$((has + h))
miss=$((miss + n - h))
offset=$((offset + n))
[ "$n" -lt 400 ] && break
done
echo "$day total=$((has + miss)) has=$has missing=$miss"
}
audit_day "$DAY"
# Sample real output:
# 20260101 total=2051 has=0 missing=2051 ← all pre-rollout
# 20260401 total=3817 has=3783 missing=34 ← post-rollout, transient gaps
# 20260424 total=3835 has=3835 missing=0 ← clean
rm -rf "$TMPDIR"To audit a range of days, just loop:
for d in 20260401 20260402 20260403; do audit_day "$d"; doneIf the user wants to know exactly which entities are missing:
LIB=6; FOLDER=14; PLUGIN=3
DAY=20260401
TMPDIR=$(mktemp -d)
> "$TMPDIR/missing.txt"
offset=0
while :; do
curl -s "http://127.0.0.1:8839/api/libraries/$LIB/folders/$FOLDER/entities?limit=400&offset=$offset&path_prefix=$HOME/.memos/screenshots/$DAY" > "$TMPDIR/batch.json"
n=$(jq 'length' "$TMPDIR/batch.json")
[ "$n" -eq 0 ] && break
jq -r ".[] | select((.plugin_status | map(.plugin_id) | index($PLUGIN)) == null) | \"\(.id) \(.filepath)\"" "$TMPDIR/batch.json" >> "$TMPDIR/missing.txt"
offset=$((offset + n))
[ "$n" -lt 400 ] && break
done
wc -l "$TMPDIR/missing.txt"
head "$TMPDIR/missing.txt"memos scan <PATH> --plugin <PLUGIN_ID> walks the directory, checks each file's plugin_status, and triggers the structured_vlm webhook only for entities missing that plugin (idempotent — already-processed entities are skipped).
# Backfill one day's structured_vlm
memos scan "$HOME/.memos/screenshots/20260401" --plugin 3 -bs 4-bs 4 (batch-size 4) is a good balance — high enough to keep the VLM endpoint busy, low enough to recover gracefully if a batch fails. The default -bs 1 is too slow for whole-day backfills.
Warn the user about cost first:
tokens × $price/1M × 3000Quote the estimate before running. Don't auto-trigger backfill of multiple days without explicit "yes, do all of them".
Re-run the audit (Step 3) for the day. missing should be near 0. If it's still high, the VLM endpoint may be unreachable or returning errors — check ~/.memos/logs/ for structured_vlm failure-category lines.
primary.what (or no structured_vlm_v1_* metadata at all): tell the user "these look like pre-rollout / failed entities; want me to backfill the day?"Don't trigger backfill silently. Always show the audit numbers + cost estimate first, get explicit user confirmation, then run.
The filepath is an absolute local path. To show the user the actual image:
# macOS
open "/Users/.../screenshots/20260423/00001.webp"
# Linux
xdg-open "/path/to/file.webp"
# Windows (PowerShell)
Start-Process "C:\path\to\file.webp"Don't open more than 2-3 at a time — overwhelming. Pick the top hit, show its timestamp + primary.what, and ask if user wants more.
q="" returns recent files unranked — that's crud.list_entities behavior, not search. If user says "find the latest X", do q=X not q="".file_type_group='image' is hardcoded server-side — the index only contains screenshots. Don't expect to find logs / docs.active_app is the OS-reported app, not the logical product. iTerm2 running Claude Code reports active_app=iTerm2. The logical product (e.g. "Claude Code") lives in structured_vlm.primary.app. So:app_names=iTerm2q=Claude+Code (FTS hits structured_vlm metadata) over app_names.✳, ⠐, etc. Don't include spinner glyphs in q. Just use the task description text.q: URL-encode Chinese / emoji properly. curl --data-urlencode 'q=...' -G ... is safest.structured_vlm_v1_* entry — fall back to ocr_result for context.ocr_result is not in search hits — it's stripped from /api/search responses to keep them small (the full payload is ~15 KB per entry × 48 hits). The other metadata (timestamp, active_app, active_window, url, structured_vlm_*) is intact. If you genuinely need the OCR text for a specific entity, fetch /api/entities/{id} directly.网 / 站 / 中 etc.) match a lot of unrelated OCR (网络 / 网站 / 网址), so FTS will surface low-relevance hits mixed with the real ones. Hybrid RRF mitigates this somewhat via vector search, but for Chinese brand / site names always cross-check hits with URL or active_app filters (see Strategy 0) — don't trust the FTS rank alone./context endpoint behaves unlike /api/search — no ocr_result stripping, naive-UTC timestamps, and a stringified structured_vlm value. See Strategy 6 for the gotchas in full and the projection recipe.app_names). After 3 variations, ask user for more context ("do you remember which app?").found > limit: see "Handling large result sets" above. Default action depends on user intent — narrow query for casual lookup, time-slice for comprehensive scans. (P1 will add an offset parameter so the time-slice recipe can retire.)memos logs.q=X + time window and let the user scrub. (To reconstruct the local timeline around a single pinpointed moment, use Strategy 6 instead.)© arkohut, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/pensieve-search of arkohut/pensieve.
Open the folder on GitHubat commit 3e55c66
Pensieve Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pensieve Search this skillarkohut/pensieve | 1.4k | — | ~8.2k | Automated safety check: Pass | Apache-2.0 | |
| Apify Trend Analysissickn33/agentic-awesome-skills | 47k | 2 repos | ~1.2k | Automated safety check: Notes | MIT | |
| TimesFM Forecastinggoogle-research/timesfm | 34k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| StatsmodelszLanqing/codex-claude-academic-skills | 4.7k | 15 repos | ~4.9k | Automated safety check: Pass | BSD-3-Clause | |
| Timesfm ForecastingzLanqing/codex-claude-academic-skills | 4.7k | 3 repos | ~7.5k | Automated safety check: Notes | Apache-2.0 | |
| Google Maps ScraperMahanaicoach/google-maps-scraper-kit | 1.3k | — | ~2.8k | Automated safety check: Pass | MIT |
sickn33/agentic-awesome-skills
Discover and track emerging trends across Google Trends, Instagram, Facebook, YouTube, and TikTok to inform content strategy.
google-research/timesfm
Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.
zLanqing/codex-claude-academic-skills
Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.
zLanqing/codex-claude-academic-skills
Zero-shot time series forecasting with Google's TimesFM foundation model.
Mahanaicoach/google-maps-scraper-kit
Scrape Google Maps business listings (name, address, phone, website, rating, reviews, lat/lng, hours, emails) via the local gosom google-maps-scraper REST API.
timescale/pg-aiguide
A skill your agent uses to analyze an existing PostgreSQL database and identify which tables should be converted to Timescale/TimescaleDB hypertables.
Categories
Search the user's local Pensieve screenshot archive by text, app, or time range. Pensieve Search is an agent skill from arkohut/pensieve. Search the user's local Pensieve screenshot archive by text, app, or time range.
Pensieve Search fits situations like: the user asks to find a screenshot (find that thing I looked at last week; show me when I was working on the Mastra integration; find screenshots of the YouTube video about X); locate a specific moment in time across captured activity.
Run `npx skills add arkohut/pensieve --skill pensieve-search -a claude-code`. Or copy the skill folder (skills/pensieve-search in arkohut/pensieve) into .claude/skills/pensieve-search in your project. Claude Code loads it when a task matches its description.
Run `npx skills add arkohut/pensieve --skill pensieve-search -a codex`. Or copy the skill folder (skills/pensieve-search in arkohut/pensieve) into .agents/skills/pensieve-search in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add arkohut/pensieve --skill pensieve-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pensieve-search, .gemini/skills/pensieve-search, .github/skills/pensieve-search and .opencode/skills/pensieve-search in your project.
Going by SKILL.md and its folder, Pensieve Search needs the command-line tools its instructions call (jq, curl and python3).
SKILL.md names 1 domain. In commands or code: bilibili.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Pensieve Search is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.2k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Pensieve Search: Apify Trend Analysis (sickn33/agentic-awesome-skills, 47k stars), TimesFM Forecasting (google-research/timesfm, 34k stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars) and Timesfm Forecasting (zLanqing/codex-claude-academic-skills, 4.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
arkohut (a GitHub user) maintains it in arkohut/pensieve, which has 1,392 GitHub stars. The repository was last updated on July 19, 2026.
Source: arkohut/pensieve on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.