Self-hosted tool to collect and preserve webpages, media, and bookmarks in durable formats (HTML, PDF, WARC, MP4) with a CLI, web UI, and search.

28.2k stars1.6k forksMITlast commit Actively maintained
The Internet Archive Wayback Machine is a web archiving service that captures, stores, and provides access to historical snapshots of websites, enabling users to browse past versions, retrieve archived pages, and verify content changes over time.
Self-hosted tool to collect and preserve webpages, media, and bookmarks in durable formats (HTML, PDF, WARC, MP4) with a CLI, web UI, and search.

28.2k stars1.6k forksMITlast commit Actively maintained
Sosse is a Selenium-powered open-source web crawler and search engine for archiving, indexing, and monitoring dynamic websites.

410 stars25 forksAGPL-3.0last commit Actively maintained
Every option on this page is open source and free to run on your own hardware, so you own the data and there is no subscription to cancel. 2 of 2 shipped a commit in the last six months. Licences in this list: MIT, AGPL-3.0. In exchange you take on hosting, backups and updates yourself.
Browse everything in Search & Indexing Engines.