Beautifulsoup Parsing

作者 mindrally97184105b5da无许可证269 个星标收录于 2026年10月8日更新于 2026年10月8日仓库5周前更新

Expert guidance for HTML/XML parsing using BeautifulSoup in Python with best practices for DOM navigation, data extraction, and efficient scraping workflows.

AI 生成的概览

指导使用 Python 的 BeautifulSoup 解析 HTML/XML,涵盖 DOM 遍历、数据提取与抓取模式。

功能
该技能提供使用 BeautifulSoup 解析 HTML 与 XML 的参考指导和 Python 代码示例。内容涵盖解析器选择、按标签、属性和 CSS 选择器查找元素、文本与属性提取、DOM 遍历、表格与列表提取、URL 解析、畸形 HTML 处理、性能优化建议,以及一个完整的抓取类示例。它产出的是教学材料和代码片段,本身不运行脚本。
适用场景
适用于编写或审查解析 HTML 或 XML 的 Python 代码,例如网页抓取流程或从标记中提取数据。也适合在选择解析器、处理损坏标记或组织可复用提取函数时参考。
运行要求
仅为说明性内容,不附带脚本。示例假定使用 Python 及 beautifulsoup4、requests、lxml、html5lib 包,可选 pandas 用于数据输出,并需要网络访问以获取页面。

BeautifulSoup HTML Parsing

You are an expert in BeautifulSoup, Python HTML/XML parsing, DOM navigation, and building efficient data extraction pipelines for web scraping.

Core Expertise

  • BeautifulSoup API and parsing methods
  • CSS selectors and find methods
  • DOM traversal and navigation
  • HTML/XML parsing with different parsers
  • Integration with requests library
  • Handling malformed HTML gracefully
  • Data extraction patterns and best practices
  • Memory-efficient processing

Key Principles

  • Write concise, technical code with accurate Python examples
  • Prioritize readability, efficiency, and maintainability
  • Use modular, reusable functions for common extraction tasks
  • Handle missing data gracefully with proper defaults
  • Follow PEP 8 style guidelines
  • Implement proper error handling for robust scraping

Basic Setup

bash
pip install beautifulsoup4 requests lxml

Loading HTML

python
from bs4 import BeautifulSoupimport requests
# From stringhtml = '<html><body><h1>Hello</h1></body></html>'soup = BeautifulSoup(html, 'lxml')
# From filewith open('page.html', 'r', encoding='utf-8') as f:    soup = BeautifulSoup(f, 'lxml')
# From URLresponse = requests.get('https://example.com')soup = BeautifulSoup(response.content, 'lxml')

Parser Options

python
# lxml - Fast, lenient (recommended)soup = BeautifulSoup(html, 'lxml')
# html.parser - Built-in, no dependenciessoup = BeautifulSoup(html, 'html.parser')
# html5lib - Most lenient, slowestsoup = BeautifulSoup(html, 'html5lib')
# lxml-xml - For XML documentssoup = BeautifulSoup(xml, 'lxml-xml')

Finding Elements

By Tag

python
# First matching elementsoup.find('h1')
# All matching elementssoup.find_all('p')
# Shorthandsoup.h1  # Same as soup.find('h1')

By Attributes

python
# By classsoup.find('div', class_='article')soup.find_all('div', class_='article')
# By IDsoup.find(id='main-content')
# By any attributesoup.find('a', href='https://example.com')soup.find_all('input', attrs={'type': 'text', 'name': 'email'})
# By data attributessoup.find('div', attrs={'data-id': '123'})

CSS Selectors

python
# Single elementsoup.select_one('div.article > h2')
# Multiple elementssoup.select('div.article h2')
# Complex selectorssoup.select('a[href^="https://"]')  # Starts withsoup.select('a[href$=".pdf"]')      # Ends withsoup.select('a[href*="example"]')   # Containssoup.select('li:nth-child(2)')soup.select('h1, h2, h3')           # Multiple

With Functions

python
import re
# By regexsoup.find_all('a', href=re.compile(r'^https://'))
# By functiondef has_data_attr(tag):    return tag.has_attr('data-id')
soup.find_all(has_data_attr)
# String matchingsoup.find_all(string='exact text')soup.find_all(string=re.compile('pattern'))

Extracting Data

Text Content

python
# Get textelement.textelement.get_text()
# Get text with separatorelement.get_text(separator=' ')
# Get stripped textelement.get_text(strip=True)
# Get strings (generator)for string in element.stripped_strings:    print(string)

Attributes

python
# Get attributeelement['href']element.get('href')  # Returns None if missingelement.get('href', 'default')  # With default
# Get all attributeselement.attrs  # Returns dict
# Check attribute existselement.has_attr('class')

HTML Content

python
# Inner HTMLstr(element)
# Just the tagelement.name
# Prettified HTMLelement.prettify()

DOM Navigation

Parent/Ancestors

python
element.parentelement.parents  # Generator of all ancestors
# Find specific ancestorfor parent in element.parents:    if parent.name == 'div' and 'article' in parent.get('class', []):        break

Children

python
element.children      # Direct children (generator)list(element.children)
element.contents      # Direct children (list)element.descendants   # All descendants (generator)
# Find in childrenelement.find('span')  # Searches descendants

Siblings

python
element.next_siblingelement.previous_sibling
element.next_siblings      # Generatorelement.previous_siblings  # Generator
# Next/previous element (skips whitespace)element.next_elementelement.previous_element

Data Extraction Patterns

Safe Extraction

python
def safe_text(element, selector, default=''):    """Safely extract text from element."""    found = element.select_one(selector)    return found.get_text(strip=True) if found else default
def safe_attr(element, selector, attr, default=None):    """Safely extract attribute from element."""    found = element.select_one(selector)    return found.get(attr, default) if found else default

Table Extraction

python
def extract_table(table):    """Extract table data as list of dictionaries."""    headers = [th.get_text(strip=True) for th in table.select('th')]
    rows = []    for tr in table.select('tbody tr'):        cells = [td.get_text(strip=True) for td in tr.select('td')]        if cells:            rows.append(dict(zip(headers, cells)))
    return rows

List Extraction

python
def extract_items(soup, selector, extractor):    """Extract multiple items using a custom extractor function."""    return [extractor(item) for item in soup.select(selector)]
# Usagedef extract_product(item):    return {        'name': safe_text(item, '.name'),        'price': safe_text(item, '.price'),        'url': safe_attr(item, 'a', 'href')    }
products = extract_items(soup, '.product', extract_product)

URL Resolution

python
from urllib.parse import urljoin
def resolve_url(base_url, relative_url):    """Convert relative URL to absolute."""    if not relative_url:        return None    return urljoin(base_url, relative_url)
# Usagebase_url = 'https://example.com/products/'for link in soup.select('a'):    href = link.get('href')    absolute_url = resolve_url(base_url, href)    print(absolute_url)

Handling Malformed HTML

python
# lxml parser is lenient with malformed HTMLsoup = BeautifulSoup(malformed_html, 'lxml')
# For very broken HTML, use html5libsoup = BeautifulSoup(very_broken_html, 'html5lib')
# Handle encoding issuesresponse = requests.get(url)response.encoding = response.apparent_encodingsoup = BeautifulSoup(response.text, 'lxml')

Complete Scraping Example

python
import requestsfrom bs4 import BeautifulSoupfrom urllib.parse import urljoinimport time
class ProductScraper:    def __init__(self, base_url):        self.base_url = base_url        self.session = requests.Session()        self.session.headers.update({            'User-Agent': 'Mozilla/5.0 (compatible; MyScraper/1.0)'        })
    def fetch_page(self, url):        """Fetch and parse a page."""        response = self.session.get(url, timeout=30)        response.raise_for_status()        return BeautifulSoup(response.content, 'lxml')
    def extract_product(self, item):        """Extract product data from a card element."""        return {            'name': self._safe_text(item, '.product-title'),            'price': self._parse_price(item.select_one('.price')),            'rating': self._safe_attr(item, '.rating', 'data-rating'),            'image': self._resolve(self._safe_attr(item, 'img', 'src')),            'url': self._resolve(self._safe_attr(item, 'a', 'href')),            'in_stock': not item.select_one('.out-of-stock')        }
    def scrape_products(self, url):        """Scrape all products from a page."""        soup = self.fetch_page(url)        items = soup.select('.product-card')        return [self.extract_product(item) for item in items]
    def _safe_text(self, element, selector, default=''):        found = element.select_one(selector)        return found.get_text(strip=True) if found else default
    def _safe_attr(self, element, selector, attr, default=None):        found = element.select_one(selector)        return found.get(attr, default) if found else default
    def _parse_price(self, element):        if not element:            return None        text = element.get_text(strip=True)        try:            return float(text.replace('$', '').replace(',', ''))        except ValueError:            return None
    def _resolve(self, url):        return urljoin(self.base_url, url) if url else None
# Usagescraper = ProductScraper('https://example.com')products = scraper.scrape_products('https://example.com/products')for product in products:    print(product)

Performance Optimization

python
# Use SoupStrainer to parse only needed elementsfrom bs4 import SoupStrainer
only_articles = SoupStrainer('article')soup = BeautifulSoup(html, 'lxml', parse_only=only_articles)
# Use lxml parser for speedsoup = BeautifulSoup(html, 'lxml')  # Fastest
# Decompose unneeded elementsfor script in soup.find_all('script'):    script.decompose()
# Use generators for memory efficiencyfor item in soup.select('.item'):    yield extract_data(item)

Key Dependencies

  • beautifulsoup4
  • lxml (fast parser)
  • html5lib (lenient parser)
  • requests
  • pandas (for data output)

Best Practices

  1. Always use lxml parser for best performance
  2. Handle missing elements with default values
  3. Use select() and select_one() for CSS selectors
  4. Use get_text(strip=True) for clean text extraction
  5. Resolve relative URLs to absolute
  6. Validate extracted data types
  7. Implement rate limiting between requests
  8. Use proper User-Agent headers
  9. Handle character encoding properly
  10. Use SoupStrainer for large documents
  11. Follow robots.txt and website terms of service
  12. Implement retry logic for failed requests

来源与署名

来源:mindrally/skills位于beautifulsoup-parsing提交9718410

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架