Puppeteer Automation

Mindrally/skills/puppeteer-automation

作者 Mindrally7682ca77710e0971eab4e0ae5dddfa281aea0ba5无许可证269 个星标收录于 2026年10月8日更新于 2026年10月9日仓库5周前更新

Expert guidance for browser automation using Puppeteer with best practices for web scraping, testing, screenshot capture, and JavaScript execution in headless Chrome.

AI 生成的概览

提供使用 Puppeteer 编写 Node.js 浏览器自动化脚本的指导,涵盖抓取、测试和截图。

功能
该技能提供在 Node.js 中使用 Puppeteer 进行浏览器自动化的参考指导和代码模式。内容涵盖启动 Chrome 或 Chromium、页面导航、元素选择、页面交互、等待策略、截图与 PDF 生成、网络拦截、身份验证与 Cookie、浏览器上下文、错误处理以及性能优化。它产出的是示例代码片段和最佳实践建议,而非可直接运行的脚本。
适用场景
适用于构建或维护用于网页抓取、浏览器测试、截图捕获或 PDF 生成的 Puppeteer 自动化时。当你需要处理动态内容等待、拦截网络请求或管理浏览器与页面生命周期的模式时也很有用。
运行要求
需要 Node.js 和 puppeteer 包(puppeteer-core、puppeteer-cluster、puppeteer-extra 及隐身插件被列为可选依赖)。需要网络访问以连接目标网站并安装依赖包。仅包含说明文档,不含脚本或资源文件。

Puppeteer Browser Automation

You are an expert in Puppeteer, Node.js browser automation, web scraping, and building reliable automation scripts for Chrome and Chromium browsers.

Core Expertise

  • Puppeteer API and browser automation patterns
  • Page navigation and interaction
  • Element selection and manipulation
  • Screenshot and PDF generation
  • Network request interception
  • Headless and headful browser modes
  • Performance optimization and memory management
  • Integration with testing frameworks (Jest, Mocha)

Key Principles

  • Write clean, async/await based code for readability
  • Use proper error handling with try/catch blocks
  • Implement robust waiting strategies for dynamic content
  • Close browser instances properly to prevent memory leaks
  • Follow modular design patterns for reusable automation code
  • Handle browser context and page lifecycle appropriately

Project Setup

bash
npm init -ynpm install puppeteer

Basic Structure

javascript
const puppeteer = require('puppeteer');
async function main() {  const browser = await puppeteer.launch({    headless: 'new',    args: ['--no-sandbox', '--disable-setuid-sandbox']  });
  try {    const page = await browser.newPage();    await page.goto('https://example.com');    // Your automation code here  } finally {    await browser.close();  }}
main().catch(console.error);

Browser Launch Options

javascript
const browser = await puppeteer.launch({  headless: 'new',  // 'new' for new headless mode, false for visible browser  slowMo: 50,       // Slow down operations for debugging  devtools: true,   // Open DevTools automatically  args: [    '--no-sandbox',    '--disable-setuid-sandbox',    '--disable-dev-shm-usage',    '--disable-accelerated-2d-canvas',    '--disable-gpu',    '--window-size=1920,1080'  ],  defaultViewport: {    width: 1920,    height: 1080  }});

Page Navigation

javascript
// Navigate to URLawait page.goto('https://example.com', {  waitUntil: 'networkidle2',  // Wait until network is idle  timeout: 30000});
// Wait options:// - 'load': Wait for load event// - 'domcontentloaded': Wait for DOMContentLoaded event// - 'networkidle0': No network connections for 500ms// - 'networkidle2': No more than 2 network connections for 500ms
// Navigate back/forwardawait page.goBack();await page.goForward();
// Reload pageawait page.reload({ waitUntil: 'networkidle2' });

Element Selection

Query Selectors

javascript
// Single elementconst element = await page.$('selector');
// Multiple elementsconst elements = await page.$$('selector');
// Wait for elementconst element = await page.waitForSelector('selector', {  visible: true,  timeout: 5000});
// XPath selectionconst elements = await page.$x('//xpath/expression');

Evaluation in Page Context

javascript
// Get text contentconst text = await page.$eval('selector', el => el.textContent);
// Get attributeconst href = await page.$eval('a', el => el.getAttribute('href'));
// Multiple elementsconst texts = await page.$$eval('.items', elements =>  elements.map(el => el.textContent));
// Execute arbitrary JavaScriptconst result = await page.evaluate(() => {  return document.title;});

Page Interactions

Clicking

javascript
await page.click('button#submit');
// Click with optionsawait page.click('button', {  button: 'left',  // 'left', 'right', 'middle'  clickCount: 1,  delay: 100       // Time between mousedown and mouseup});
// Click and wait for navigationawait Promise.all([  page.waitForNavigation(),  page.click('a.nav-link')]);

Typing

javascript
// Type textawait page.type('input#username', 'myuser', { delay: 50 });
// Clear and typeawait page.click('input#username', { clickCount: 3 });await page.type('input#username', 'newvalue');
// Press keysawait page.keyboard.press('Enter');await page.keyboard.down('Shift');await page.keyboard.press('Tab');await page.keyboard.up('Shift');

Form Handling

javascript
// Select dropdownawait page.select('select#country', 'us');
// Check checkboxawait page.click('input[type="checkbox"]');
// File uploadconst inputFile = await page.$('input[type="file"]');await inputFile.uploadFile('/path/to/file.pdf');

Waiting Strategies

javascript
// Wait for selectorawait page.waitForSelector('.loaded');
// Wait for selector to disappearawait page.waitForSelector('.loading', { hidden: true });
// Wait for functionawait page.waitForFunction(  () => document.querySelector('.count').textContent === '10');
// Wait for navigationawait page.waitForNavigation({ waitUntil: 'networkidle2' });
// Wait for network requestawait page.waitForRequest(request =>  request.url().includes('/api/data'));
// Wait for network responseawait page.waitForResponse(response =>  response.url().includes('/api/data') && response.status() === 200);
// Fixed timeout (use sparingly)await page.waitForTimeout(1000);

Screenshots and PDFs

Screenshots

javascript
// Full page screenshotawait page.screenshot({  path: 'screenshot.png',  fullPage: true});
// Element screenshotconst element = await page.$('.chart');await element.screenshot({ path: 'chart.png' });
// Screenshot optionsawait page.screenshot({  path: 'screenshot.png',  type: 'png',  // 'png' or 'jpeg'  quality: 80,   // jpeg only, 0-100  clip: {    x: 0,    y: 0,    width: 800,    height: 600  }});

PDF Generation

javascript
await page.pdf({  path: 'document.pdf',  format: 'A4',  printBackground: true,  margin: {    top: '20px',    right: '20px',    bottom: '20px',    left: '20px'  }});

Network Interception

javascript
// Enable request interceptionawait page.setRequestInterception(true);
page.on('request', request => {  // Block images and stylesheets  if (['image', 'stylesheet'].includes(request.resourceType())) {    request.abort();  } else {    request.continue();  }});
// Modify requestspage.on('request', request => {  request.continue({    headers: {      ...request.headers(),      'X-Custom-Header': 'value'    }  });});
// Monitor responsespage.on('response', async response => {  if (response.url().includes('/api/')) {    const data = await response.json();    console.log('API Response:', data);  }});

Authentication and Cookies

javascript
// Basic HTTP authenticationawait page.authenticate({  username: 'user',  password: 'pass'});
// Set cookiesawait page.setCookie({  name: 'session',  value: 'abc123',  domain: 'example.com'});
// Get cookiesconst cookies = await page.cookies();
// Clear cookiesawait page.deleteCookie({ name: 'session' });

Browser Context and Multiple Pages

javascript
// Create incognito contextconst context = await browser.createIncognitoBrowserContext();const page = await context.newPage();
// Multiple pagesconst page1 = await browser.newPage();const page2 = await browser.newPage();
// Get all pagesconst pages = await browser.pages();
// Handle popupspage.on('popup', async popup => {  await popup.waitForLoadState();  console.log('Popup URL:', popup.url());});

Error Handling

javascript
async function scrapeWithRetry(url, maxRetries = 3) {  for (let i = 0; i < maxRetries; i++) {    try {      const browser = await puppeteer.launch();      const page = await browser.newPage();
      // Set timeout      page.setDefaultTimeout(30000);
      await page.goto(url, { waitUntil: 'networkidle2' });      const data = await page.$eval('.content', el => el.textContent);
      await browser.close();      return data;    } catch (error) {      console.error(`Attempt ${i + 1} failed:`, error.message);      if (i === maxRetries - 1) throw error;      await new Promise(r => setTimeout(r, 2000 * (i + 1)));    }  }}

Performance Optimization

javascript
// Disable unnecessary featuresawait page.setRequestInterception(true);page.on('request', request => {  const blockedTypes = ['image', 'stylesheet', 'font'];  if (blockedTypes.includes(request.resourceType())) {    request.abort();  } else {    request.continue();  }});
// Reuse browser instanceconst browser = await puppeteer.launch();
async function scrape(url) {  const page = await browser.newPage();  try {    await page.goto(url);    // ... scraping logic  } finally {    await page.close();  // Close page, not browser  }}
// Use connection pool for parallel scrapingconst cluster = require('puppeteer-cluster');

Key Dependencies

  • puppeteer
  • puppeteer-core (for custom Chrome installations)
  • puppeteer-cluster (for parallel scraping)
  • puppeteer-extra (for plugins)
  • puppeteer-extra-plugin-stealth (anti-detection)

Best Practices

  1. Always close browser instances in finally blocks
  2. Use waitForSelector before interacting with elements
  3. Prefer networkidle2 over networkidle0 for faster loads
  4. Use stealth plugin for anti-bot bypass
  5. Implement proper error handling and retries
  6. Monitor memory usage in long-running scripts
  7. Use browser context for isolated sessions
  8. Set reasonable timeouts for all operations

来源与署名

来源:Mindrally/skills位于puppeteer-automation提交7682ca7

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架