Agent skill

Opensearch Function Scoring Algorithms

by pproenca in pproenca/dot-skills

Search relevance and ranking on OpenSearch/Elasticsearch for a two-sided marketplace — candidate retrieval (hybrid BM25 + kNN, RRF, two-tower EBR), base relevance (BM25F, multimatch, LambdaMART)…

MITAuto-check passedAI & LLM Engineering

Install Opensearch Function Scoring Algorithms

skills CLI
$ npx skills add pproenca/dot-skills --skill opensearch-function-scoring-algorithms -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pproenca/dot-skills opensearch-function-scoring-algorithms --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/opensearch-function-scoring-algorithms .claude/skills/opensearch-function-scoring-algorithms && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opensearch-function-scoring-algorithms
GitHub stars
215
Token cost
~3.8k tokens
SKILL.md length
1,215 words
Files
61 (incl. references, assets)
Skills in repo
41
Repo updated
First seen
Licence
MIT

At a glance

Search relevance and ranking on OpenSearch/Elasticsearch for a two-sided marketplace — candidate retrieval (hybrid BM25 + kNN, RRF, two-tower EBR), base relevance (BM25F, multimatch, LambdaMART)…

  • Works in 9 steps: Candidate Retrieval & Recall (CRITICAL) → Base Relevance & Field Scoring (CRITICAL) → Quality Signals & Confidence Bounds (HIGH) → …
  • Learning-to-rank
  • SKILL.md covers When to Apply, The Search Ranking Lifecycle, Rule Categories by Priority and Quick Reference, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Opensearch Function Scoring Algorithms is an agent skill from pproenca/dot-skills. Search relevance and ranking on OpenSearch/Elasticsearch for a two-sided marketplace — candidate retrieval (hybrid BM25 + kNN, RRF, two-tower EBR), base relevance (BM25F, multimatch, LambdaMART), quality signals (Wilson lower bound, Bayesian average, rankfeature saturation/sigmoid), personalization (listing/user/session embeddings), spatial/temporal decay (gauss/exp), marketplace balance (conversion-weighted ranking, supply fairness, Pareto multi-objective), bias correction (IPS, click models, Thompson sampling)…

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 63 other files, including reference files and assets (for example `AGENTS.md`, `assets/templates/_template.md` and `metadata.json`).

It sits in AI & LLM Engineering, covering Search implementation, Embeddings and A/B testing. It works with OpenSearch and Elasticsearch. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.

When your agent uses it

  • Learning-to-rank
  • Two-sided ranking
  • Exposure fairness
  • Judgment set construction

Example prompts

  • “why is my OpenSearch ranking bad”
  • “/opensearch-function-scoring-algorithms”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Candidate Retrieval & Recall (CRITICAL)
  2. Base Relevance & Field Scoring (CRITICAL)
  3. Quality Signals & Confidence Bounds (HIGH)
  4. Personalization & Embeddings (HIGH)
  5. Spatial & Temporal Decay (HIGH)
  6. Two-Sided Marketplace Balance (HIGH)
  7. Bias Correction & Online Learning (HIGH)
  8. Evaluation & Measurement (HIGH)
  9. Diversity & Re-ranking (MEDIUM-HIGH)

What it can do on your machine

Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Opensearch Function Scoring Algorithms loads about 3.8k tokens when it runs, and up to ~55k if it reads all its reference files. Until then it costs about 242 tokens; SKILL.md has 1,215 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~242
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~55k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 1,215 words, ~3,798 tokens.

Download SKILL.mdSave it as .claude/skills/opensearch-function-scoring-algorithms/SKILL.md (or your agent's skills folder). This skill also uses 60 other files; get the full folder from GitHub.
name
opensearch-function-scoring-algorithms
description
Search relevance and ranking on OpenSearch/Elasticsearch for a two-sided marketplace — candidate retrieval (hybrid BM25 + kNN, RRF, two-tower EBR), base relevance (BM25F, multi_match, LambdaMART), quality signals (Wilson lower bound, Bayesian average, rank_feature saturation/sigmoid), personalization (listing/user/session embeddings), spatial/temporal decay (gauss/exp), marketplace balance (conversion-weighted ranking, supply fairness, Pareto multi-objective), bias correction (IPS, click models, Thompson sampling), empirical evaluation (judgment sets, NDCG, ablation, A/B sizing, CUPED, regression suites), and diversity (MMR, DPP, max-per-host). Triggers on function_score, rank_feature, script_score, kNN, hybrid query, learning-to-rank, two-sided ranking, exposure fairness, NDCG, A/B testing, judgment set construction, ranking ablation, or "why is my OpenSearch ranking bad". Applies to Elasticsearch too — same APIs.

Marketplace-Research OpenSearch Function Scoring Best Practices

A reference distillation of research-backed algorithms for ranking in two-sided marketplaces (Airbnb, Uber Eats, DoorDash, Etsy, eBay, Booking.com) implemented on OpenSearch or Elasticsearch. Contains 56 rules across 9 categories, prioritised by cascade effect in the search ranking pipeline. Each rule explains the WHY (the cascade or the bias it corrects), shows incorrect-vs-correct code (OpenSearch JSON queries, Painless scripts, Python pre-processing, evaluation methodology), and links to the canonical source — KDD/SIGIR/WSDM papers, the OpenSearch documentation, and the engineering blogs of the marketplaces that proved these patterns at scale.

When to Apply

Reach for this skill when:

  • Designing a new marketplace search system on OpenSearch or Elasticsearch from scratch
  • Tuning function_score / rank_feature / script_score queries that aren't moving the needle
  • Setting up hybrid retrieval (BM25 + dense vectors) with Reciprocal Rank Fusion
  • Choosing between HNSW and IVF for billion-scale ANN indexes
  • Adding personalization via listing/user embeddings or two-tower architectures
  • Correcting position bias in click logs before retraining an LTR model
  • Designing exposure-fairness or new-listing cold-start exposure allocation
  • Composing decay functions (gauss / exp / linear) over geo + date + freshness
  • Diversifying the top window with MMR, DPP, or per-host caps
  • Debugging "why does my top-10 show 8 listings from one host?" or "why does ranking favor popular incumbents?"
  • Building offline evaluation infrastructure — graded judgment sets, NDCG@k pipelines, ablation studies, regression query suites
  • Designing A/B tests for ranking changes — MDE / power / sample-size pre-computation, CUPED variance reduction, online-offline correlation calibration
  • Attributing lift to specific scoring components — "did my new bias-correction help, or was it the embeddings, or both?"

The rules apply to any OpenSearch/Elasticsearch-backed marketplace search regardless of vertical — accommodation, food delivery, restaurants, services, jobs, secondhand goods, real estate. Triggers include "marketplace ranking", "search relevance", "function_score", "rank_feature", "script_score", "kNN", "hybrid search", "RRF", "learning to rank", "embedding-based retrieval", "two-tower", "position bias", "MMR", "supply fairness", "Pareto multi-objective", "NDCG", "judgment set", "ablation study", "CUPED", "A/B sample size", "ranking eval", and "why are my search results bad".

The Search Ranking Lifecycle

Categories are derived from the marketplace search ranking pipeline. Earlier stages cascade — a miss in recall (stage 1) cannot be repaired by any downstream boost, and a wrong base relevance multiplies through every functional score:

text
Query → [1] Recall → [2] Base Relevance → [3] Quality Signals → [4] Personalization
      → [5] Geo/Time Decay → [6] Marketplace Balance → [7] Diversity Re-rank → Results
                                                            ↑
                                          [8] Bias Correction (applied across all stages
                                                       and into training)
                                                            ↑
                                          [9] Evaluation & Measurement (the meta-layer:
                                                       judgment sets, NDCG, ablation, A/B
                                                       sizing, CUPED — without these you
                                                       can't tell if any rule helped)

Rule Categories by Priority

PriorityCategoryImpactPrefixRules
1Candidate Retrieval & RecallCRITICALrecall-6
2Base Relevance & Field ScoringCRITICALrel-7
3Quality Signals & Confidence BoundsHIGHqual-6
4Personalization & EmbeddingsHIGHpers-7
5Spatial & Temporal DecayHIGHdecay-5
6Two-Sided Marketplace BalanceHIGHmarket-7
7Bias Correction & Online LearningHIGHbias-6
8Evaluation & MeasurementHIGHeval-7
9Diversity & Re-rankingMEDIUM-HIGHdiv-5

Quick Reference

1. Candidate Retrieval & Recall (CRITICAL)
2. Base Relevance & Field Scoring (CRITICAL)
3. Quality Signals & Confidence Bounds (HIGH)
4. Personalization & Embeddings (HIGH)
5. Spatial & Temporal Decay (HIGH)
Show full SKILL.md (487 more words)Show less
6. Two-Sided Marketplace Balance (HIGH)
7. Bias Correction & Online Learning (HIGH)
8. Evaluation & Measurement (HIGH)
9. Diversity & Re-ranking (MEDIUM-HIGH)

How to Use

For a focused question ("which decay function for geo distance?"), jump directly to the relevant rule (decay-gauss-geo) — each rule is self-contained with the WHY, OpenSearch query/Painless code, and the canonical source citation.

For a full ranking system review, work the categories top-to-bottom. The cascade ordering is real: get recall right first (no boost recovers a missed candidate), then base relevance (it's the multiplicand of every functional score), then quality / personalization / decay / marketplace balance / bias correction in that order. Diversity is the last re-rank step over a well-ordered top window.

For correcting bias before retraining, start with bias-position-ips and bias-click-models — applying IPS to position-confounded click data is the single highest-leverage change for any marketplace that retrains LTR models on logged clicks.

For testing multiple algorithms together and validating empirically, start with eval-graded-judgment-set (build the foundation), eval-ndcg-primary-metric (pick the metric), then eval-ablation-attribution (attribute lift to specific components). Pair with eval-online-offline-correlation to verify your offline metric predicts online behavior, eval-ab-sample-size-mde + eval-cuped-variance-reduction for disciplined A/B testing, and eval-regression-query-suite to catch silent quality drops on named queries.

For research-citing a design decision, every rule ends with the canonical reference — KDD/SIGIR/WSDM papers, the relevant engineering blog (Airbnb, Pinterest, DoorDash, Etsy, Just Eat Takeaway, Thumbtack), or the OpenSearch documentation page.

Read section definitions for the cascade-impact rationale behind the category ordering, or the rule template when adding a new rule.

Reference Files

FileDescription
references/_sections.mdCategory definitions and ordering by cascade impact
AGENTS.mdCompact TOC navigation (auto-built; do not edit by hand)
assets/templates/_template.mdTemplate for authoring new rules
metadata.jsonVersion and authoritative reference URLs

© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 60 other files (references, assets) in skills/.experimental/opensearch-function-scoring-algorithms of pproenca/dot-skills.

  • SKILL.md
  • AGENTS.md
  • assets/templates/_template.md
  • metadata.json
  • references/_sections.md
  • references/bias-click-models.md
  • references/bias-counterfactual-eval.md
  • references/bias-interleaved-evaluation.md
  • references/bias-popularity-debiasing.md
  • references/bias-position-ips.md
  • references/bias-thompson-sampling.md
  • references/decay-exp-freshness.md
  • references/decay-gauss-geo.md
  • references/decay-multi-field-composition.md
  • references/decay-offset-noise.md
  • references/decay-scale-calibration.md
  • references/div-category-diversity.md
  • references/div-dpp-quality-diversity.md
  • … and 43 more

Open the folder on GitHubat commit cf93c57

Compare with similar skills

Opensearch Function Scoring Algorithms next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opensearch Function Scoring Algorithms compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opensearch Function Scoring Algorithms this skillpproenca/dot-skills215—~3.8kAutomated safety check: PassMIT
Amazon Opensearch Serviceaws/agent-toolkit-for-aws2.8k—~2.4kAutomated safety check: PassApache-2.0
Docker Compose Testsjillesvangurp/kt-search155—~295Automated safety check: PassMIT
Detecting Insider Threat With Uebamukul975/Anthropic-Cybersecurity-Skills34k—~738Automated safety check: PassApache-2.0
Analytics Opensearch Expertiseaws/tools-for-devops-agent100—~6.9kAutomated safety check: PassApache-2.0
Hybrid Search Implementationwshobson/agents40k10 repos~497Automated safety check: PassMIT

Similar skills

  • Amazon Opensearch Service

    aws/agent-toolkit-for-aws

    Official

    Guides migration, provisioning, search, log-analytics, trace-analytics, and Agentic AI Assistant workflows for Amazon OpenSearch Service and Serverless across six capabilities — migration…

    2.8k GitHub stars~2.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Docker Compose Tests

    jillesvangurp/kt-search

    Use Gradle Compose tasks to prepare and recover local Elasticsearch/OpenSearch test infrastructure in kt-search.

    155 GitHub stars~295 tokensUpdated 3 mo ago
    MobileAuto-check passed
  • Detecting Insider Threat With Ueba

    mukul975/Anthropic-Cybersecurity-Skills

    Implement User and Entity Behavior Analytics (UEBA) using Elasticsearch/OpenSearch to build behavioral baselines, calculate anomaly scores, perform peer group analysis, and alert on insider threat…

    34k GitHub stars~738 tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Analytics Opensearch Expertise

    aws/tools-for-devops-agent

    Official

    Amazon OpenSearch Service domain health assessment. An agent skill from aws/tools-for-devops-agent.

    100 GitHub stars~6.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Shows how to run vector and keyword search side by side and merge their results, so retrieval catches both meaning and exact terms in RAG and search systems.

    40k GitHub starsUsed in 10 repos~497 tokens
    AI & LLM EngineeringAuto-check passed
  • Storing And Querying Vectors

    aws/agent-toolkit-for-aws

    Official

    Store and query vector embeddings using Amazon S3 Vectors, a cost-effective long-term vector storage service with its own API namespace (s3vectors).

    2.8k GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from pproenca/dot-skills

All 41 skills in this repo
  • Audio Voice Recovery

    pproenca/dot-skills

    Audio forensics and voice recovery guidelines for CSI-level audio analysis.

    215 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Codemod React Pipeline

    pproenca/dot-skills

    Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.

    215 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dev Rfc

    pproenca/dot-skills

    Create well-structured RFCs and technical proposals for software projects.

    215 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Dx Harness

    pproenca/dot-skills

    Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.

    215 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Language Spec Author

    pproenca/dot-skills

    Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…

    215 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    215 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Opensearch Function Scoring Algorithms

What does Opensearch Function Scoring Algorithms do?

Search relevance and ranking on OpenSearch/Elasticsearch for a two-sided marketplace — candidate retrieval (hybrid BM25 + kNN, RRF, two-tower EBR), base relevance (BM25F, multimatch, LambdaMART)…. Opensearch Function Scoring Algorithms is an agent skill from pproenca/dot-skills.

When should I use Opensearch Function Scoring Algorithms?

Opensearch Function Scoring Algorithms fits situations like: learning-to-rank; two-sided ranking; exposure fairness; judgment set construction.

How do I install Opensearch Function Scoring Algorithms in Claude Code?

Run `npx skills add pproenca/dot-skills --skill opensearch-function-scoring-algorithms -a claude-code`. Or copy the skill folder (skills/.experimental/opensearch-function-scoring-algorithms in pproenca/dot-skills) into .claude/skills/opensearch-function-scoring-algorithms in your project. Claude Code loads it when a task matches its description.

How do I install Opensearch Function Scoring Algorithms in Codex?

Run `npx skills add pproenca/dot-skills --skill opensearch-function-scoring-algorithms -a codex`. Or copy the skill folder (skills/.experimental/opensearch-function-scoring-algorithms in pproenca/dot-skills) into .agents/skills/opensearch-function-scoring-algorithms in your project. Codex loads it when a task matches its description.

Can I use Opensearch Function Scoring Algorithms in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill opensearch-function-scoring-algorithms -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opensearch-function-scoring-algorithms, .gemini/skills/opensearch-function-scoring-algorithms, .github/skills/opensearch-function-scoring-algorithms and .opencode/skills/opensearch-function-scoring-algorithms in your project.

What does Opensearch Function Scoring Algorithms need to run?

SKILL.md names no scripts, command-line tools or credentials: Opensearch Function Scoring Algorithms is instructions for the agent only. Our summary lists: Python 3.

Does Opensearch Function Scoring Algorithms access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Opensearch Function Scoring Algorithms safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Opensearch Function Scoring Algorithms use?

Opensearch Function Scoring Algorithms is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opensearch Function Scoring Algorithms use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 52k tokens, read only when the agent opens those files.

What are the alternatives to Opensearch Function Scoring Algorithms?

Skills that share tags, products or a category with Opensearch Function Scoring Algorithms: Amazon Opensearch Service (aws/agent-toolkit-for-aws, 2.8k stars), Docker Compose Tests (jillesvangurp/kt-search, 155 stars), Detecting Insider Threat With Ueba (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Analytics Opensearch Expertise (aws/tools-for-devops-agent, 100 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opensearch Function Scoring Algorithms?

pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 215 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on August 15, 2026.

Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.