1
0
Fork 0
docsearch/docs/src/inside-the-engine.md
Sylvain Pace a94f4c57dc
Doc/update content (#434)
* # This is a combination of 2 commits.
# This is the 1st commit message:

# This is a combination of 3 commits.
# This is the 1st commit message:

# This is a combination of 3 commits.
# This is the 1st commit message:

chore(deps): update dependency onchange to v4.1.0

integrate previous work

enhance tyle/content

reformat part 1

wait for review

adance start_urls

enhance attributes description

fix typo

proofread documentation/docsearch

add apiKey mention

intefrate review and small fixes

finished proofreading

update sclient-rendering

use unseen review

# This is the commit message #2:

update README #269

# This is the commit message #3:

fix json

# This is the commit message #2:

enhance  as algolia/docsearch-configs#387

# This is the commit message #3:

updating flavicon

# This is the commit message #2:

Update 1-customize-configuration-file.html.md.erb

* # This is a combination of 2 commits.
# This is the 1st commit message:

# This is a combination of 3 commits.
# This is the 1st commit message:

# This is a combination of 3 commits.
# This is the 1st commit message:

chore(deps): update dependency onchange to v4.1.0

integrate previous work

enhance tyle/content

reformat part 1

wait for review

adance start_urls

enhance attributes description

fix typo

proofread documentation/docsearch

add apiKey mention

intefrate review and small fixes

finished proofreading

update sclient-rendering

use unseen review

# This is the commit message #2:

update README #269

# This is the commit message #3:

fix json

# This is the commit message #2:

enhance  as algolia/docsearch-configs#387

# This is the commit message #3:

updating flavicon

# This is the commit message #2:

Update 1-customize-configuration-file.html.md.erb

* documenting algolia/docsearch-scraper#387

* documenting algolia/docsearch-scraper#387

* fixing changelog

* fixing typo

* # This is a combination of 2 commits.
# This is the 1st commit message:

# This is a combination of 3 commits.
# This is the 1st commit message:

# This is a combination of 3 commits.
# This is the 1st commit message:

chore(deps): update dependency onchange to v4.1.0

integrate previous work

enhance tyle/content

reformat part 1

wait for review

adance start_urls

enhance attributes description

fix typo

proofread documentation/docsearch

add apiKey mention

intefrate review and small fixes

finished proofreading

update sclient-rendering

use unseen review

# This is the commit message #2:

update README #269

# This is the commit message #3:

fix json

# This is the commit message #2:

enhance  as algolia/docsearch-configs#387

# This is the commit message #3:

updating flavicon

# This is the commit message #2:

Update 1-customize-configuration-file.html.md.erb

* # This is a combination of 2 commits.
# This is the 1st commit message:

# This is a combination of 3 commits.
# This is the 1st commit message:

# This is a combination of 3 commits.
# This is the 1st commit message:

chore(deps): update dependency onchange to v4.1.0

integrate previous work

enhance tyle/content

reformat part 1

wait for review

adance start_urls

enhance attributes description

fix typo

proofread documentation/docsearch

add apiKey mention

intefrate review and small fixes

finished proofreading

update sclient-rendering

use unseen review

# This is the commit message #2:

update README #269

# This is the commit message #3:

fix json

# This is the commit message #2:

enhance  as algolia/docsearch-configs#387

# This is the commit message #3:

updating flavicon

# This is the commit message #2:

Update 1-customize-configuration-file.html.md.erb

* documenting algolia/docsearch-scraper#387

* documenting algolia/docsearch-scraper#387

* fixing typo

* fixing style issue

* fixing recurrnet issue

* general github handle

* fixing missed conflict

* fix and closes algolia/docsearch-configs-private#82

* remove builded content

* adding hitsperpage

* fix

* enhance

* doc proof reading part 1

* adding how to build an index and updating content

* reformat

* textlint

* typos

* Update behavior.md

* Update config-file.md

* Update crawler-overview.md

* Update dropdown.md

* Update faq.md

* Update how-do-we-build-an-index.md

* Update how-does-it-work.md

* Update inside-the-engine.md

* Update integrations.md

* Update styling.md

* Update tips.md

* Update who-can-apply.md

* Update integrations.md

* Usefull but sometime we do want to ignore it on purpose

* 4fca279b9a

* auto fix & formating
2018-09-04 08:47:38 +02:00

68 lines
3 KiB
Markdown

---
layout: two-columns
title: Inside the engine
---
This page will explain in more detail how the crawler extracts content from your
page every 24h, and how it ranks the results.
## Crawling
Each crawl will begin its journey by the value of the `start_urls` you have in
your config. It will read those pages and recursively extract and follow every
link in those pages until it has browsed every compliant page.
If you have explicitly defined a `sitemap.xml`, our crawler will scrap every
provided and compliant page. We do recommend using [a sitemap][1] since it
clearly exposes URLs to crawl and avoid missing pages that aren't linked from
another one.
## Extracting content
Building records using the scraper is pretty intuitive. According to your
settings, we extract the payload of your web page and index it, preserving your
data's structure. This is achieved in a simple way:
- We **read top down** your web page following your HTML flow and pick out your
matching elements according to their **levels** based on the `selectors_level`
defined.
- We create a record for each paragraph along with its hierarchical path. This
construction is based on their **time of appearance** along the flow.
- We **index** these records with the appropriate global settings (e.g.
metadata, tags, etc.)
_**Note:** The above process performs sanity tests as it scrapes, in order to
detect errors. If indeed there are any serious warnings, it will abort and hence
not overwrite your current index. These checks ensure that your dedicated index
isn't flushed._
## Ranking records
Algolia always returns the most relevant results first, using a [tie-breaking
approach][2]. DocSearch will first search for exact matches in your keywords
then fallback to partial matches. Those results will then be ordered based, once
again, on the page hierarchy, as extracted from the `selectors`.
The default strategy is to promote records having matching words in the highest
level fist. Thus if two results have the same matching words, the one having
them in the highest level (lvl0) will be reanked higher. We also use the
position of the matching words. The sooner they appear within the HTML flow, the
higher the record will be ranked.
The relevancy is based on several factors and can be customized according to the
Algolia tie-breaking method.
You can boost pages depending ont their URLs. This is done from the `start_urls`
and its `page_rank` attributes. It is a numeric value, default to 0. The higher
it is, the higher results from the matching pages will be ranked. For example
all pages with a `page_rank` of 5 will be returned before pages with a
`page_rank` of 1.
You could even change the relevancy strategy by [overwriting the default
`customRanking`][3] used by the index by using the `custom_settings` option of
your config.
[1]: https://www.sitemaps.org/
[2]:
https://www.algolia.com/doc/guides/ranking/ranking-formula/#tie-breaking-approach
[3]: https://www.algolia.com/doc/guides/ranking/custom-ranking/