Skip to content
EntityQ3097891· pop 7· linked from 43 articles

Heritrix is a web crawler designed for web archiving. It was originally written in collaboration between the Internet Archive, National Library of Norway and National Library of Iceland. Heritrix is available under a free software license and written in Java. The main interface is accessible using a web browser, and there is a command-line tool that can optionally be used to initiate crawls.

Described at

Heritrix is web crawler used by the Internet Archive, which provides a web-based user interface after initial configuration on a Linux machine. Also used by the Library of Congress, Heritrix captures metadata in the Web ARChive (WARC) format.

Excerpt from a page describing this subject · 2,140 chars · not written by Vinony

Source code

Heritrix is designed to respect the robots.txt exclusion directives and META nofollow tags. Please consider the load your crawl will place on seed sites and set politeness policies accordingly. Also, always identify your crawl with contact information in the User-Agent so sites that may be adversely affected by your crawl can contact you or adapt their server behavior accordingly. Some individual source code files are subject to or offered under other licenses. See the included LICENSE.txt file for more information.

Excerpt from the source-code README · 3,111 chars · not written by Vinony

Wikidata facts

Official website
heritrix.readthedocs.io
Image
Heritrix-screenshot.png
Show 5 more facts
software version identifier
3.14.1
source code repository URL
github.com/internetarchive/heritrix3
Commons category
Heritrix
Sources (8)

via Wikidata · CC0

~5 min read

Article

9 sections
Contents
  • Projects using Heritrix
  • Arc files
  • Tools for processing Arc files
  • Command-line tools
  • See also
  • References
  • External links
  • Tools by Internet Archive
  • Links to related tools

Heritrix is a web crawler designed for web archiving. It was originally written in collaboration between the Internet Archive, National Library of Norway and National Library of Iceland. Heritrix is available under a free software license and written in Java. The main interface is accessible using a web browser, and there is a command-line tool that can optionally be used to initiate crawls.

Heritrix was developed jointly by the Internet Archive and the Nordic national libraries on specifications written in early 2003. The first official release was in January 2004, and it has been continually improved by employees of the Internet Archive and other interested parties.

Available in 6 languages

via Wikidata sitelinks · CC0

Connections

Categories