1
0
Fork 0
docsearch/packages/website/docs/record-extractor.md
Paul Jankowski ecd905d440
feat: promote DocSearch v5 to main (#2968)
* feat(askai): add compatibility with algolia mcp search tool [DASH-2294] (#2862)

## Summary
Fixes DASH-2294
Add compatibility with the Algolia MCP search tool (`algolia_search_index_${string}`) in AskAI.

## Changes
- Add `AlgoliaMCPSearchTool` type to handle the Algolia MCP server search tool
- Refactor how number of hits are retrieved in `ToolCall` into a `getNumberOfHits` helper

## Test plan
- Added unit tests for modified code 

* chore: Update to use tsdown for build system (#2824)

* chore: Update to use tsdown for build system

* fix: docsearch-react build

* fix: lint

* fix: glob resolved to incorrect version

* chore: migrate to from yarn, lerna and shipjs to bun & changesets (#2827)

* chore: add tool-versions file for node and bun versions (#2866)

* chore: watch in parallel (#2867)

* feat: agent studio feedback integration (#2868)

* feat(askai): Split Ask AI modal into own component (#2884)

* feat(askai): Split Ask AI modal into own component

* refactor(react): share modal utilities

* refactor(react): share search box form

* refactor(react): extract start screen sections

* refactor(react): extract shared modal hooks

* fix: lint adapter

* refactor(react): reorganize modal files

* fix: type error in examples

* fix: remove ai modal from adapter for now, fix import paths of react package

* feat(askai): Agent Studio core tools (#2886)

* feat(askai): Implement dynamic tool calls

* move ToolCall to components dir

* Converge Agent Studio search tools to same definition, fix client side tools breaking UI state

* add examples for custom tools

* fix: lint & types

* feat(askai): add Agent Studio memory support (#2888)

* feat(askai): remove Ask AI transport layer (#2889)

* feat(askai): add Agent Studio memory support

* refactor(askai): remove Ask AI transport abstraction

* feat(askai): Feedback notes and tags (#2890)

* feat(askai): add Agent Studio memory support

* refactor(askai): remove Ask AI transport abstraction

* feat(askai): Feedback notes and tags

* fix: bump css bundle size limit

* move feedback actions to components

* chore: fix deploys for v5 branch

* feat(askai): Aggregate MCP search tool calls (#2891)

* feat(askai): add Agent Studio memory support

* refactor(askai): remove Ask AI transport abstraction

* feat(askai): Feedback notes and tags

* fix: bump css bundle size limit

* move feedback actions to components

* feat(askai): Aggregate MCP search tool calls

* feat(askai): Allow dynamic indices for Agent Studio (#2893)

* feat(v5): UI updates (#2896)

* feat(v5): UI updates

* fix: css file size

* fix: e2e tests

* fix: e2e tests

* fix: e2e tests

* chore: add theme toggle to react demo example

* Update sources panel display, update dark theme

* fix: lint

* fix(askai): address ui review feedback

* fix: pin icon positioning

* fix(askai): improve a11y and dark-mode shimmer for thinking and error states

- add role=alert/status and aria-hidden on error/thinking UI
- support dark-mode shimmer gradients via CSS variables
- respect prefers-reduced-motion for shimmer
- handle null date in useRelativeFormattedDate with fallback translation

* feat(v5): Add hit breadcrumbs (#2897)

* feat(v5): UI updates

* fix: css file size

* fix: e2e tests

* fix: e2e tests

* fix: e2e tests

* chore: add theme toggle to react demo example

* Update sources panel display, update dark theme

* fix: lint

* fix(askai): address ui review feedback

* fix: pin icon positioning

* fix(askai): improve a11y and dark-mode shimmer for thinking and error states

- add role=alert/status and aria-hidden on error/thinking UI
- support dark-mode shimmer gradients via CSS variables
- respect prefers-reduced-motion for shimmer
- handle null date in useRelativeFormattedDate with fallback translation

* feat(v5): Add hit breadcrumbs

* fix: bump css bundle size limit

* Fix after conflicts

* chore: move CSS building to lightning css (#2898)

* feat: Facet filters for search (#2899)

* feat(v5): Initial facet filters work

* Perf updates, dark theme, facet chips, a11y improvements

* fix: bump css bundle size limit

* Dedupe facet filters, refetch facets on searchParameters changes

* Add chevron flourish

* fix(askai): Fix new conversation causing thread depth errors (#2900)

* feat(v5): Add hit result badge (#2901)

* feat(v5): Add hit result badge

* Add background to hit result badge

* feat(v5): Add follow up prompt suggestions (#2902)

* feat(v5): Add follow up prompt suggestions

* fix: bump css bundle size limit

* docs(agents): document Cursor Cloud dev environment setup for v5 (Bun) (#2903)

Co-authored-by: Cursor Agent <cursoragent@cursor.com>

* feat(mcp): setup mcp plugins (#2895)

* feat(v5): Add prompt suggestions to keyword search (#2912)

* feat(v5): Add prompt suggestions to keyword search

* cleanup: Move consistent object to reusable constant

* chore(v5): Split Ask AI related CSS into own bundle (#2913)

* chore(v5): Split Ask AI related CSS into own bundle

* move style.css to include modal and askai

* fix: Ensure stage level and watch level scripts use bun runtime (#2915)

* fix: Ensure stage level and watch level scripts use bun runtime

* chore: move to node@24 update imports

* fix: lint

* feat(js): Document JS based hybrid mode, fix JS packages (#2916)

* feat(js): Document JS based hybrid mode, fix JS packages

* update: add model onOpen to docs

* feat(cli): add @docsearch/cli for MCP setup and search (#2911)

* chore(tsdown): Bump to latest tsdown version (#2918)

* chore(tsdown): Bump to latest tsdown version

* fix: bump nvmrc node version

* fix: cli tsconfig

* fix: website build

* fix: example build

* fix: circleci install bun

* fix: lint

* fix: circleci install bun

* fix: circleci install bun

* fix: circleci install bun

* refactor(docusaurus-adapter): rework theme config for v5 and modularize SearchPage (#2904)

Co-authored-by: Paul Jankowski <8BitTitan@gmail.com>

* feat(askai): Move askai related props under root askai (#2919)

* feat(askai): Move askai related props under root askai

* fix: playwright test case

* fix(docusaurus): validate Ask AI options

* feat(js): Split JS bundles for search only (#2920)

* chore: Move to oxlint and oxfmt (#2923)

* chore: Get NPM OIDC token before publishing (#2924)

* chore: Enter v5 beta (#2925)

* chore: Enter pre release mode for v5

* chore: update release summary

* chore: version bump

* fix: Remove NPM_ID_TOKEN for release

* fix: Try setting blank NPM_TOKEN

* fix: Try blank NPM_AUTH_TOKEN

* chore: bump node and npm for release job

* docs(mcp): add service disclaimer (#2921)

* fix: Agent Studio MCP search tool (#2927)

* fix: Agent Studio MCP search tool

* Add changeset

* chore: Update stylelint (#2926)

* chore: stylelint update

* bun.lock

* Add changeset

* fix(website): use bare @import for tailwindcss (#2933)

Tailwind's build-time `@import` cannot be written with `url()` notation,
so `@import url('tailwindcss')` was passed through as a plain CSS import
instead of being processed by Tailwind.

Also syncs bun.lock with the 5.0.0-beta.0 versions already committed to
package.json.

* chore: version v5.0.0-beta.1 (#2932)

* feat(react): remove deprecated index props (#2936)

* feat(mcp): add ChatGPT and Codex DocSearch plugin package (#2938)

* fix: cleanup claude

* feat: website redesign (#2930)

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: lint

* fix: crash on the demo (#2940)

* feat(docs): Document v5 beta (#2935)

* chore(docs): v5 documentation

* Writing style clean up

* fix: website after conflicts

* fix: reported issues on mobile (#2944)

* chore: Introduce a11y smoke tests (#2943)

* chore: Add Lorris as codeowner (#2946)

* fix(askai): Ask AI fixes for v5 (#2945)

* fix: General v5 fixes (#2947)

- Fix `ref` console error for a `FacetMenu`
- Whitespace only search/conversation input does not trigger requests
- Fix flash of no results page on search

* feat: v5 general improvements (#2948)

* feat(v5): General fixes and improvements

* add changeset

* fix: bundlesize

* feat(v5): UI and DX improvements (#2949)

* feat: Rename assistantId to agentId

* feat: Allow reading default facet values from index searchParameters

* feat: Remove indexName prop from Sidepanel, cleanup documentation pages

* feat: Move appId and apiKey up into @docsearch/core

* feat: Add back nested grouping of search results

* add changeset

* revert changes to example demo

* fix: e2e tests

* chore: push git tags on version release (#2951)

* chore: release v5.0.0-beta.2 (#2950)

* chore: Fix pushing git tags (#2953)

* fix(askai): sanitize markdown HTML in v5 (#2954)

Backport of #2929.\n\nOriginal commit: 681cbfec03

Co-authored-by: Vasco Bettencourt <32492444+vascobettencourt@users.noreply.github.com>

* fix(v5): stop truncating mobile snippets (#2958)

* fix(v5): stop truncating mobile snippets

Backport of #2907.\n\nOriginal commit: 9ad6d169fe

* fix(v5): allow mobile hit text to wrap

Completes the v5 adaptation of #2907 by overriding later v5 child-level truncation rules.\n\nOriginal commit: 9ad6d169fe

* chore(v5): account for mobile wrapping CSS

Updates the CSS size budget for the v5 adaptation of #2907.\n\nOriginal commit: 9ad6d169fe

---------

Co-authored-by: Divyansh Singh <40380293+brc-dd@users.noreply.github.com>

* feat(v5): Add new footerAction prop (#2952)

* feat(v5): Add new footerAction prop

* Resolve PR comments

* fix(v5): recognize conversation depth errors (#2957)

Backport of #2881.\n\nOriginal commit: f68e52251c

Co-authored-by: Felipe Bermudez <felipeberm@gmail.com>

* fix(v5): expose Sidepanel search parameter types (#2956)

* fix(v5): expose Sidepanel search parameter types

Backport of #2906.\n\nOriginal commit: 4710d0ca77

* Delete sidepanel.test.ts

Had a pointless test case in it.

---------

Co-authored-by: Divyansh Singh <40380293+brc-dd@users.noreply.github.com>

* fix(v5): ignore slash shortcut on focused buttons (#2955)

Backport of #2871.\n\nOriginal commit: 0e41a78c44

Co-authored-by: Sigmabro <122412346+Sigmabrogz@users.noreply.github.com>

* fix(agentStudio): agents dynamic mode enabled (#2959)

* fix(agentStudio): agents dynamic mode enabled

* fix(askai): use string[] for dynamic agentStudio indices

* feat(docs): add Ask AI to Agent Studio migration guide (#2931)

* feat(docs): add Ask AI to Agent Studio migration guide

* feat(docs): agentStudio migrating from askAI

* feat(docs): renaming agentId

* feat(agentStudio): dynamic mode indices updated

* fix: Docusaurus adapter styling, DocSearch website fixes (#2960)

* chore: release v5.0.0-beta.3 (#2961)

* feat(website): Launch updates (#2964)

* fix(website): Fix font loading (#2966)

* feat(v5): Back port cost control errors (#2965)

* chore: release v5.0.0-beta.4 (#2967)

---------

Co-authored-by: Vincent Lemeunier <vincentlemeunier+git@gmail.com>
Co-authored-by: Dylan Tientcheu <dylan.tientcheu@algolia.com>
Co-authored-by: Lorris Saint-Genez <lorrissaintgenez@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Dylan Tientcheu <dylantientcheu@gmail.com>
Co-authored-by: Vasco Bettencourt <32492444+vascobettencourt@users.noreply.github.com>
Co-authored-by: Divyansh Singh <40380293+brc-dd@users.noreply.github.com>
Co-authored-by: Felipe Bermudez <felipeberm@gmail.com>
Co-authored-by: Sigmabro <122412346+Sigmabrogz@users.noreply.github.com>
2026-08-06 15:31:45 -04:00

11 KiB

title description
Record Extractor Configure the DocSearch record extractor for Algolia Crawler records.

Introduction

:::info

This page documents the helpers.docsearch method. See the Algolia Crawler documentation for information about the Algolia Crawler.

:::

Set the recordExtractor parameter on an action to extract each page. Its function returns the data to index as an array of JSON objects.

The helpers are functions for extracting content and generating Algolia records.

Usage

The most common way to use the DocSearch helper is to return its result to the recordExtractor function.

recordExtractor: ({ helpers }) => {
  return helpers.docsearch({
    recordProps: {
      lvl0: {
        selectors: "header h1",
      },
      lvl1: "article h2",
      lvl2: "article h3",
      lvl3: "article h4",
      lvl4: "article h5",
      lvl5: "article h6",
      content: "main p, main li",
    },
  });
},

Manipulate the DOM with Cheerio

The Cheerio instance ($) allows you to manipulate the DOM:

recordExtractor: ({ $, helpers }) => {
  // Removing DOM elements we don't want to crawl
  $(".my-warning-message").remove();

  return helpers.docsearch({
    recordProps: {
      lvl0: {
        selectors: "header h1",
      },
      lvl1: "article h2",
      lvl2: "article h3",
      lvl3: "article h4",
      lvl4: "article h5",
      lvl5: "article h6",
      content: "main p, main li",
    },
  });
},

Provide fallback selectors

Fallback selectors can be useful when retrieving content that might not exist in some pages:

recordExtractor: ({ $, helpers }) => {
  return helpers.docsearch({
    recordProps: {
      // `.exists h1` will be selected if `.exists-probably h1` does not exists.
      lvl0: {
        selectors: [".exists-probably h1", ".exists h1"],
      },
      lvl1: "article h2",
      lvl2: "article h3",
      lvl3: "article h4",
      lvl4: "article h5",
      lvl5: "article h6",
      // `.exists p, .exists li` will be selected.
      content: [
        ".does-not-exists p, .does-not-exists li",
        ".exists p, .exists li",
      ],
    },
  });
},

Provide raw text (defaultValue)

Only the lvl0 and custom variables selectors support this option

You might want to structure your search results differently than your website, or provide a defaultValue to a potentially non-existent selector:

recordExtractor: ({ $, helpers }) => {
  return helpers.docsearch({
    recordProps: {
      lvl0: {
        // It also supports the fallback DOM selectors syntax!
        selectors: ".exists-probably h1",
        defaultValue: "myRawTextIfDoesNotExists",
      },
      lvl1: "article h2",
      lvl2: "article h3",
      lvl3: "article h4",
      lvl4: "article h5",
      lvl5: "article h6",
      content: "main p, main li",
      // The variables below can be used to filter your search
      language: {
        // It also supports the fallback DOM selectors syntax!
        selectors: ".exists-probably .language",
        // Since custom variables are used for filtering, we allow sending
        // multiple raw values
        defaultValue: ["en", "en-US"],
      },
    },
  });
},

Indexing content for faceting

These selectors also support defaultValue and fallback selectors

To index content for frontend filters, such as version or language, define custom variables in recordProps. The helper adds them to each matching Algolia record:

recordExtractor: ({ helpers }) => {
  return helpers.docsearch({
    recordProps: {
      lvl0: {
        selectors: "header h1",
      },
      lvl1: "article h2",
      lvl2: "article h3",
      lvl3: "article h4",
      lvl4: "article h5",
      lvl5: "article h6",
      content: "main p, main li",
      // The variables below can be used to filter your search
      foo: ".bar",
      language: {
        // It also supports the fallback DOM selectors syntax!
        selectors: ".does-not-exists",
        // Since custom variables are used for filtering, we allow sending
        // multiple raw values
        defaultValue: ["en", "en-US"],
      },
      version: {
        // You can send raw values without `selectors`
        defaultValue: ["latest", "stable"],
      },
    },
  });
},

The version, language, and foo attributes are then available in your records:

{
  "foo": "valueFromBarSelector",
  "language": ["en", "en-US"],
  "version": ["latest", "stable"]
}

Add every filter attribute to the index's attributesForFaceting, then expose up to five of them with the v5 facets option. If you display one with resultBadgeKey, also add that attribute to attributesToRetrieve; see the resultBadgeKey reference.

V5 result breadcrumbs use the hierarchy.lvl0 through hierarchy.lvl6 values generated from your selectors. Keep the heading levels ordered and include the hierarchy attributes in attributesToRetrieve.

Boost search results with pageRank

This parameter allows you to boost records using a custom ranking attribute built from the current pathsToMatch. Pages with highest pageRank will be returned before pages with a lower pageRank. The default value is 0 and you can pass any numeric value as a string, including negative values.

Search results are sorted by weight (desc), so you can have both boosted and non boosted results. The weight of each result will be computed for a given query based on multiple factors: match level, position, etc. and the pageRank value will be added to this final weight. The pageRank on its own may not be enough to influence the results of your query depending on how your overall ranking is set up. If changing the pageRank value doesn't influence your search results enough, even with large values, move weight.pageRank higher in the Ranking and Sorting page for your index.

You can view the computed weight directly from the Algolia dashboard (dashboard.algolia.com->search->perform a search->mouse hover over the "ranking criteria" icon bottom right of each record). That will give you an idea of what pageRank value is acceptable for your case.

{
  indexName: "YOUR_INDEX_NAME",
  pathsToMatch: ["https://YOUR_WEBSITE_URL/api/**"],
  recordExtractor: ({ $, helpers, url }) => {
    const isDocPage = /\/[\w-]+\/docs\//.test(url.pathname);
    const isBlogPage = /\/[\w-]+\/blog\//.test(url.pathname);
    return helpers.docsearch({
      recordProps: {
        lvl0: {
          selectors: "header h1",
        },
        lvl1: "article h2",
        lvl2: "article h3",
        lvl3: "article h4",
        lvl4: "article h5",
        lvl5: "article h6",
        content: "article p, article li",
        pageRank: isDocPage ? "-2000" : isBlogPage ? "-1000" : "0",
      },
    });
  },
},

Reduce the number of records

If you encounter the Extractors returned too many records error when your page outputs more than 750 records, the aggregateContent option helps you reduce the number of records at the content level of the extractor.

{
  indexName: "YOUR_INDEX_NAME",
  pathsToMatch: ["https://YOUR_WEBSITE_URL/api/**"],
  recordExtractor: ({ $, helpers }) => {
    return helpers.docsearch({
      recordProps: {
        lvl0: {
          selectors: "header h1",
        },
        lvl1: "article h2",
        lvl2: "article h3",
        lvl3: "article h4",
        lvl4: "article h5",
        lvl5: "article h6",
        content: "article p, article li",
      },
      aggregateContent: true,
    });
  },
},

Reduce the record size

If you encounter the Records extracted are too big error, your records or source page might contain too much information. The recordVersion option reduces record size by removing fields used only by the DocSearch v2 UI.

{
  indexName: "YOUR_INDEX_NAME",
  pathsToMatch: ["https://YOUR_WEBSITE_URL/api/**"],
  recordExtractor: ({ $, helpers }) => {
    return helpers.docsearch({
      recordProps: {
        lvl0: {
          selectors: "header h1",
        },
        lvl1: "article h2",
        lvl2: "article h3",
        lvl3: "article h4",
        lvl4: "article h5",
        lvl5: "article h6",
        content: "article p, article li",
      },
      recordVersion: "v3",
    });
  },
},

recordProps API Reference

lvl0

type: Lvl0 | required

type Lvl0 = {
  selectors: string | string[];
  defaultValue?: string;
};

lvl1, content

type: string | string[] | required

lvl2, lvl3, lvl4, lvl5, lvl6

type: string | string[] | optional

pageRank

type: number | optional

See the live example

Custom variables

type: string | string[] | CustomVariable | optional

type CustomVariable =
  | {
      defaultValue: string | string[];
    }
  | {
      selectors: string | string[];
      defaultValue?: string | string[];
    };

Define custom variables in recordProps. You can use them with v5 facets and per-index filters.

helpers.docsearch API Reference

aggregateContent

type: boolean | default: true | optional

This option groups the Algolia records created at the content level of the selector into a single record for its matching heading.

recordVersion

type: 'v3' | 'v2' | default: v2 | optional

This option selects the crawler record schema. It doesn't select the DocSearch UI package version. Set it to v3 to remove fields used only by the DocSearch v2 UI. The v3 value is also the current record schema for DocSearch v5 frontends.

indexHeadings

type: boolean | { from: number, to: number } | default: true | optional

This option tells the crawler if the headings (lvlX) should be indexed.

  • When false, only records for the content level will be created.
  • When from, to is provided, only records for the lvlX to lvlY will be created.