1. Crawling.
1. Crawling.
The very first and most important step of the search engines is the crawling. Crawling is the process by which search engine bots read your website. The bots visit your website pages, posts, images and every other content on your website in order to read it. The crawling is done by special software called crawlers (or colloquially, spiders or bots). For example, the google crawlers identify as googlebot. If you look at your website server for all the user-agents visiting your website, you will notice one called googlebot. The way these crawlers discover pages/posts and crawl them is by following URLs. Say for example the crawler comes to your homepage. Then on the homepage you have a menu-bar that contains all the important links of your website. The crawlers will follow all those links on your menu and read them. By doing so, that's how the crawlers are able to index almost the entire internet. The main goal of crawling the internet is to find new pages and update the existing ones. Hence you as an SEO expert, if you want to get all your website and web pages indexed, you need to make them available to the crawlers. You need to create a sitemap.xml file with all your links and submit it to the search engine. You cannot rank on search engines like google without first getting your website and webpages indexed.
2. Indexing.
2. Indexing.
Once the search engines crawl the internet, the next step is to index the URLs that they have crawled. You know when we talk about the internet, we are talking about tens of trillions of URLs. The search engines surely need a system to organize all these trillions of URLs. The search engines need to know how to organize the links based on value. Moreover, the search engines need to know which link is which. The engines need to know what link goes where. What each link is all about. Is it a link for a page, or a blog-post or an image or a file or what? The search engine systems need to do all this data analysis. All this data from all the links that the search engine has crawled is analyzed and organized into a massive database called index. Hence the term search engine indexing. But now i know you are asking what is stored? The search engines do not store everything. The reason for this is because there is a lot of junk on the internet. Hence the search engines have to filter a lot of this junk and leave it out. Some of the content that is stored is information like keywords and freshness(how recent is your content). In addition to that, the search engines also store the page structure of things like headings and link structure. Moreover, the engines also store things like media, you know images, videos and files. Now i know you have a new question, why is indexing necessary? The indexing is necessary so that the search engines can quickly retrieve relevant results when a user searches. If it weren't for indexing, search engines would have to go and crawl the page every time one searches anything.
3. Ranking (Retrieval).
3. Ranking (Retrieval).
Now this is the final step of how search engines work. This is the step that we all get to interact with the search engines. Every time you search on any search engine, you are on this step of the search engines. When you type any keywords on the search engine, the search engine checks its index (from step 2 above) to find the most relevant results. The engine checks its index to find the most relevant pages based on hundreds of factors called ranking signals. In this article we are not going to talk about the ranking signals. If you want to learn more about the search engines ranking signals we can make a separate article for that. Just to give you a slight insights of the ranking signals, here is a small list.- Relevance to the search query (does the page match the keywords?).
- Page authority (how trustworthy is the site?).
- User experience (is the page fast and mobile-friendly?).
- Freshness of content.
- Number and quality of backlinks.
