Curate a new Web scraping collection - #5283
Conversation
Maintainer triageCollection
|
| Item | Stars | Last push | Owner type | Notes |
|---|---|---|---|---|
scrapy/scrapy |
63,536 | 2026-08-01 | Organization | – |
apify/crawlee |
25,143 | 2026-08-01 | Organization | – |
apify/crawlee-python |
9,385 | 2026-07-30 | Organization | – |
gocolly/colly |
25,402 | 2026-06-18 | Organization | – |
D4Vinci/Scrapling |
72,090 | 2026-07-30 | User | – |
spider-rs/spider |
2,634 | 2026-08-01 | Organization | – |
projectdiscovery/katana |
17,253 | 2026-07-28 | Organization | – |
firecrawl/firecrawl |
159,001 | 2026-08-01 | Organization | – |
unclecode/crawl4ai |
75,772 | 2026-07-30 | User | – |
ScrapeGraphAI/Scrapegraph-ai |
28,862 | 2026-07-20 | Organization | – |
microsoft/playwright |
93,797 | 2026-07-31 | Organization | – |
puppeteer/puppeteer |
95,390 | 2026-08-01 | Organization | – |
SeleniumHQ/selenium |
34,344 | 2026-08-01 | Organization | – |
seleniumbase/SeleniumBase |
12,909 | 2026-07-26 | Organization | – |
lightpanda-io/browser |
33,278 | 2026-08-01 | Organization | – |
browserless/browserless |
13,542 | 2026-07-31 | Organization | – |
cheeriojs/cheerio |
30,430 | 2026-07-31 | Organization | – |
jhy/jsoup |
11,379 | 2026-08-01 | User | – |
sparklemotion/nokogiri |
6,275 | 2026-07-27 | Organization | – |
PuerkitoBio/goquery |
14,976 | 2026-07-31 | Organization | – |
symfony/panther |
3,064 | 2026-06-04 | Organization | – |
dgtlmoon/changedetection.io |
32,574 | 2026-07-31 | User | – |
getmaxun/maxun |
16,968 | 2026-07-31 | Organization | – |
crawlee-cloud/crawlee-cloud |
34 | 2026-07-30 | Organization | – |
b0cefe6 to
8ff89bd
Compare
tomthorogood
left a comment
There was a problem hiding this comment.
Thank you for your contribution. One small request before we can approve this change.
| display_name: Web scraping | ||
| created_by: aminembarki | ||
| --- | ||
| The web is the world's largest dataset, and these open source projects help developers collect it: crawling frameworks for Python, JavaScript, Go, Rust, Java, Ruby, and PHP, headless browser automation, HTML parsers, and self-hosted platforms for running crawlers at scale. |
There was a problem hiding this comment.
crawling frameworks for Python, JavaScript, Go, Rust, Java, Ruby, and PHP, headless browser automation, HTML parsers, and self-hosted platforms for running crawlers at scale.
I don't think it's helpful to list the various integrations/languages here; as this topic grows, this may no longer stay accurate, so it's better to let the projects and frameworks speak for themselves.
There was a problem hiding this comment.
Fair point. I've simplified it to:
The web is the world's largest dataset. These open source projects help developers crawl, extract, and run scrapers at scale.
Just pushed the change.
| - symfony/panther | ||
| - dgtlmoon/changedetection.io | ||
| - getmaxun/maxun | ||
| - crawlee-cloud/crawlee-cloud |
There was a problem hiding this comment.
I appreciate you being up front about your ownership of this product. It is a little blurry, re: our self-promotion clause, but I had a look through the project's contributions and history and don't think there's any ill intent meant by including it here, so we can allow it this time.
There was a problem hiding this comment.
Thanks a lot for looking into it 🙏
Please confirm this pull request meets the following requirements:
Which change are you proposing?
Curating a new topic or collection
https://github.com/topics/[NAME]orhttps://github.com/collections/[NAME])*.pngimage (if applicable) andindex.mdindex.mdconform to the Style Guide and API docs: https://github.com/github/explore/tree/main/docsWeb scraping is one of the most active areas of open source — the
web-scrapingtopic alone has tens of thousands of repositories, and projects like Scrapy, Firecrawl, Crawlee, and Playwright are among the most-starred on GitHub — yet Explore has no collection for it (the closest,opensource-testinganddigital-preservation, cover browser automation only for QA and archiving). This collection curates the ecosystem end to end: crawling frameworks across seven languages (Python, JavaScript, Go, Rust, Java, Ruby, and PHP), AI-era extraction tools, headless browser automation, HTML parsers, and self-hosted platforms. Every entry was checked to be actively maintained and non-archived as of July 2026; historically important but archived projects (PhantomJS, pyspider, Portia) were deliberately excluded.For transparency: I maintain one of the smaller entries (crawlee-cloud/crawlee-cloud) and have ordered it last. Although it is young on GitHub, it is not a toy project — it runs in production today, orchestrating roughly 190 scraper runs daily. That said, I'm happy to drop it from the list if the maintainers prefer.