1
0
Fork 0
docsearch/docs/src/crawler-overview.md
Sylvain Pace 25f7214830
Docs/update (#641)
* fixing typo

* keeping same convention

* avoid wrong use of word crawling

* removing superfluous comment

* make the documentation more explicit to prevent error in the key usage, introduces feedback from algolia/docsearch-scraper#438, give instruction about how to use the headless chrome

* add synonym, resolves #640

* precise sitemap

* fix typos

* removing unavailable anchor

* fixing wrong URL

* resolves #518 by removing unwanted char from anchors

* fixing linter issues

* fix remark-lint warning
2019-03-28 14:32:31 +01:00

35 lines
1.3 KiB
Markdown

---
layout: two-columns
title: Crawler Overview
---
## How?
The DocSearch crawler is written in python and heavily based on the [Scrapy][1]
framework. It will crawl all pages of your website and extract content from the
HTML structure to populate an Algolia index.
It will automatically follow every internal links to make sure we are not
missing any content, and will use the semantics of your HTML structure to
construct its records. This means that `h1`,`h2`, etc., (`selectors`) titles
will be used for the hierarchy, and each `p` of text will be used as a potential
result.
Those CSS selectors can be overwritten, and each website actually has its own
JSON configuration file that describes in more detail how the crawl should
behave. You can find the complete list of options in [the related section][2].
## When?
We automatically run each config every 24h. This is done from our own
infrastructure, meaning that you don't need to install anything on your side. We
run this service entirely free of charge, but we're asking that you keep the
"powered by Algolia" logo next to the search results.
That being said, if you'd like to [run DocSearch on your own][3], all the code
is open source and even packaged as a Docker image. Download it, and run it with
your own credentials.
[1]: https://scrapy.org/
[2]: ./config-file.html
[3]: ./run-your-own.html