Web Crawler diagram template

A polite distributed crawler with a URL frontier, fetchers, parsers and dedup storage.

Web Crawler architecture diagramOpen in ArchBoard

Builds a new scene in your browser. Your existing scenes are not touched.

About this design

A crawler is a loop: take a URL, fetch it, extract links, and feed the new ones back in. Doing that at scale means the URL frontier becomes the heart of the system. It orders work by priority and enforces politeness, ensuring one host is never hammered by many fetchers at once. Fetcher workers resolve DNS, honour robots rules, download pages and place raw content in object storage. Parsers pull out text and links, a deduplication step using content hashes and a bloom filter keeps the crawler from revisiting the same page, and fresh links return to the frontier. A scheduler decides recrawl frequency from how often a page changes. Discuss traps such as infinite calendars, handling JavaScript-rendered pages, and splitting the frontier by host hash so each worker owns a stable set of domains.

Diagram as text

This is the source of the diagram, in the ArchBoard diagram DSL. Paste it into Tools, Diagram from text to rebuild or change it.

title "Web crawler"
direction LR
scheduler "Seed scheduler" -> queue kafka "URL frontier"
url-frontier -> worker fetch "Fetchers"
[fetch x4]
fetch -> external "Websites"
fetch -> storage s3 "Raw pages"
raw-pages -> worker parse "Parsers"
parse -> cache redis "Seen URLs"
parse -> url-frontier
parse -> search elasticsearch "Index"

More interview classics templates