The context API to search, scrape, and interact with the web at scale. 🔥
Scraper GitHub Repositories
Explore popular GitHub repositories tagged “scraper”.
Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.
Trending Repositories
Create agents that monitor and act on your behalf. Your agents are standing by!
A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。
👾 Fast and simple video download library and CLI tool written in Go
The fast, flexible, and elegant library for parsing and manipulating HTML and XML.
Open source AI job application bot in Python: auto apply to jobs, with a tailored resume and cover letter for each posting.
Elegant Scraper and Crawler Framework for Golang
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
🚀「Douyin_TikTok_Download_API」是一个开箱即用的高性能异步抖音、快手、TikTok、Bilibili数据爬取工具,支持API调用,在线批量解析及下载。
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
:orange_book: 中华新华字典数据库。包括歇后语,成语,词语,汉字。
AV 电影管理系统, avmoo , javbus , javlibrary 爬虫,线上 AV 影片图书馆,AV 磁力链接数据库,Japanese Adult Video Library,Adult Video Magnet Links - Japanese Adult Video Database
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
A Smart, Automatic, Fast and Lightweight Web Scraper for Python
A collection of awesome web crawler,spider in different languages
A Chrome DevTools Protocol driver for web automation and scraping.
Turn any webpage into structured data using LLMs
Distributed crawler powered by Headless Chrome
A social networking service scraper in Python
💡 Download the complete source code of any website (including all assets). [ Javascripts, Stylesheets, Images ] using Node.js
Analysis of Bot Protection systems with available countermeasures 🚿. How to defeat anti-bot system 👻 and get around browser fingerprinting scripts 🕵️♂️ when scraping the web?
Swiss-army tool for scraping and extracting data from online assets, made for hackers
YouTube video downloader in javascript.
Twitter API Scraper | Without an API key | Twitter Internal API | Free | Twitter scraper | Twitter Bot
A library that scrapes Linkedin for user data
A community-driven way to read and chat with AI bots - powered by chatGPT.
Scrape all the media from an OnlyFans account - Updated regularly
🔮 A Node.js scraper for humans.
Emby/Jellyfin 的一个日本电影刮削器插件,可以从某些网站抓取影片信息。