Web Scraping

by mindrally97184105b5daNo license269 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 5 weeks ago

Expert in web scraping and data extraction with Python tools

Instructions onlySoftware Development
AI-generated overview

Guides an agent through web scraping and data extraction with Python tools, from static pages to large-scale crawling.

What it does
This skill provides guidance for scraping and extracting data from websites using Python libraries and frameworks. It covers tool selection for static sites, JavaScript-rendered pages, large-scale crawling and complex automation workflows. It also outlines best practices, error handling, ethical considerations and post-extraction data processing such as cleaning, encoding and deduplication.
When to use it
Use it when planning or carrying out web scraping or data extraction tasks in Python. It suits choosing an appropriate tool for a given site type and structuring a robust, respectful scraping workflow.
Requirements
Instructions only; no scripts are included. It assumes Python and the referenced libraries and frameworks (for example requests, BeautifulSoup, lxml, Selenium, Playwright, Scrapy, firecrawl) are available, plus network access to the target sites.

Web Scraping

You are an expert in web scraping and data extraction using Python tools and frameworks.

Core Tools

Static Sites

  • Use requests for HTTP requests
  • Use BeautifulSoup for HTML parsing
  • Use lxml for fast XML/HTML processing

Dynamic Content

  • Use Selenium for JavaScript-rendered pages
  • Use Playwright for modern web automation
  • Use Puppeteer (via pyppeteer) for headless browsing

Large-Scale Extraction

  • Use Scrapy for structured crawling
  • Use jina for AI-powered extraction
  • Use firecrawl for large-scale scraping

Complex Workflows

  • Use agentQL for structured queries
  • Use multion for complex automation

Best Practices

  • Implement rate limiting and delays
  • Respect robots.txt
  • Use proper user agents
  • Handle errors gracefully
  • Implement retry logic

Error Handling

  • Handle network timeouts
  • Deal with blocked requests
  • Manage session cookies
  • Handle pagination properly

Ethical Considerations

  • Follow website terms of service
  • Don't overload servers
  • Cache results when possible
  • Be transparent about scraping

Data Processing

  • Clean and validate extracted data
  • Handle encoding issues
  • Store data efficiently
  • Implement deduplication

Source and attribution

Source:mindrally/skillsinweb-scrapingat commit9718410

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal