Trending repositories: web-scraping

15 tracked repositories tagged with web-scraping, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

15 of 15 repositories

  • firecrawl/firecrawl

    Supercharge your AI agents with data from the web and beyond. Building the library for superintelligence. 🔥

    AI summary: An API service that crawls websites and turns them into clean, LLM-ready markdown data.

    188,588dataTypeScriptAGPL-3.0
  • browser-use/browser-use

    Agents that use the browser.

    AI summary: A robust Python library empowering AI agents to autonomously navigate and automate dynamic web browser interactions.

    117,132developer-toolsPythonMIT
  • D4Vinci/Scrapling

    🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev

    AI summary: Adaptive Python web scraping framework supporting everything from single requests to full-scale concurrent crawls.

    85,671developer-toolsPythonBSD-3-Clause
  • vercel-labs/agent-browser

    Browser automation CLI for AI agents

    AI summary: A headless browser environment specifically optimized for AI agents to navigate, scrape, and interact with the web.

    43,477developer-toolsRustApache-2.0
  • lightpanda-io/browser

    Lightpanda: the headless browser designed for AI and automation

    AI summary: A headless browser built in Zig specifically designed for AI agents and web scraping.

    35,766developer-toolsZigAGPL-3.0
  • JCodesMore/ai-website-cloner-template

    Clone any website with one command using AI coding agents

    AI summary: A Next.js boilerplate optimized for AI agents to quickly clone and rebuild target website layouts.

    35,594webTypeScriptMIT
  • CloakHQ/CloakBrowser

    Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.

    AI summary: A privacy-first web browser designed to seamlessly bypass anti-bot protections and tracking.

    31,877developer-toolsPythonMIT
  • jackwener/OpenCLI

    Make Any Website into CLI & Use your logged-in browser by AI agent.

    AI summary: A runtime that converts any website or local tool into a deterministic CLI and allows AI agents to operate logged-in browsers.

    29,771developer-toolsJavaScriptApache-2.0
  • browser-use/browser-harness

    Browser Harness | Self-healing harness that enables LLMs to complete any task.

    AI summary: A robust CDP-based bridge for connecting AI agents directly to live browser sessions for complex web automation.

    18,283developer-toolsPythonMIT
  • wechat-article/wechat-article-exporter

    一款在线的 微信公众号文章批量下载 工具,支持导出阅读量与评论数据,无需搭建任何环境,可通过 在线网站 使用,支持 docker 私有化部署和 Cloudflare 部署。 支持下载各种文件格式,其中 HTML 格式可100%还原文章排版与样式。

    AI summary: An online tool and Docker service for bulk downloading WeChat Official Account articles with perfect HTML formatting.

    13,015productivityTypeScriptMIT
  • jo-inc/camofox-browser

    Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.

    AI summary: An anti-detection browser server that provides fingerprint spoofing at the C++ level for AI agents.

    11,343securityJavaScriptMIT
  • pinchtab/pinchtab

    High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.

    AI summary: A token-efficient HTTP API for browser control designed specifically for AI agents.

    10,331developer-toolsGoMIT
  • Tencent/BrowserSkill

    Let AI agents use your real, logged-in browser without interrupting your work. CLI + extension for browser automation across any shell-capable AI agent.

    AI summary: A tool that connects AI agents directly to a user's authenticated web browser.

    8,111productivityTypeScriptMIT
  • KnockOutEZ/wigolo

    The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.

    AI summary: A keyless, privacy-first web search and extraction layer built specifically for AI agents.

    5,442developer-toolsTypeScriptOther
  • tinyfish-io/bigset-oss

    Open-source BigSet — self-hostable live datasets populated by TinyFish web agents

    AI summary: An experimental tool that uses autonomous AI agents to research and build structured datasets from the live web.

    1,702dataTypeScriptAGPL-3.0