Web Reader Skill
This skill guides the implementation of web page reading and content extraction functionality using the z-ai-web-dev-sdk package, enabling applications to fetch and process web page content programmatically.
Skills Path
Skill Location: {project_path}/skills/web-reader
This skill is located at the above path in your project.
Reference Scripts: Example test scripts are available in the {Skill Location}/scripts/ directory for quick testing and reference. See {Skill Location}/scripts/web-reader.ts for a working example.
Overview
Web Reader allows you to build applications that can extract content from web pages, retrieve article metadata, and process HTML content. The API automatically handles content extraction, providing clean, structured data from any web URL.
IMPORTANT: z-ai-web-dev-sdk MUST be used in backend code only. Never use it in client-side code.
Prerequisites
The z-ai-web-dev-sdk package is already installed. Import it as shown in the examples below.
CLI Usage (For Simple Tasks)
For simple web page content extraction, you can use the z-ai CLI instead of writing code. This is ideal for quick content scraping, testing URLs, or simple automation tasks.
Basic Page Reading
Save Page Content
Common Use Cases
CLI Parameters
--name, -n: Required - Function name (use "page_reader")--args, -a: Required - JSON arguments object with:url(string, required): The URL of the web page to read
--output, -o <path>: Optional - Output file path (JSON format)
Response Structure
The CLI returns a JSON object containing:
title: Page titlehtml: Main content HTMLtext: Plain text contentpublish_time: Publication timestamp (if available)url: Original URLmetadata: Additional page metadata
Example Response
Processing Multiple URLs
When to Use CLI vs SDK
Use CLI for:
- Quick content extraction
- Testing URL accessibility
- Simple web scraping tasks
- One-off content retrieval
Use SDK for:
- Batch URL processing with custom logic
- Integration with web applications
- Complex content processing pipelines
- Production applications with error handling
How It Works
The Web Reader uses the page_reader function to:
- Fetch the web page content
- Extract main article content and metadata
- Parse and clean the HTML
- Return structured data including title, content, and publication time
Basic Web Reading Implementation
Simple Page Reading
Extract Article Text Only
Read Multiple Pages
Advanced Use Cases
Web Content Analyzer
RSS Feed Reader
Content Aggregator
Web Scraping Pipeline
Response Format
Successful Response
Response Fields
Best Practices
1. Error Handling
2. Rate Limiting
3. Caching Strategy
4. Parallel Processing
5. Content Processing
Common Use Cases
- News Aggregation: Collect and aggregate news articles from multiple sources
- Content Monitoring: Track changes on specific web pages
- Research Tools: Extract information from academic or reference websites
- Price Tracking: Monitor product pages for price changes
- SEO Analysis: Extract page metadata and content for SEO purposes
- Archive Creation: Create local copies of web content
- Content Curation: Collect and organize web content by topic
- Competitive Intelligence: Monitor competitor websites for updates
Integration Examples
Express.js API Endpoint
Scheduled Content Fetcher
Troubleshooting
Issue: "SDK must be used in backend"
- Solution: Ensure z-ai-web-dev-sdk is only imported and used in server-side code
Issue: Failed to fetch page (404, 403, etc.)
- Solution: Verify the URL is accessible and not behind authentication/paywall
Issue: Incomplete or missing content
- Solution: Some pages may have dynamic content that requires JavaScript. The reader extracts static HTML content.
Issue: High token usage
- Solution: The token usage depends on page size. Consider caching frequently accessed pages.
Issue: Slow response times
- Solution: Implement caching, use parallel processing for multiple URLs, and consider rate limiting
Issue: Empty HTML content
- Solution: Check if the page requires authentication or has anti-scraping measures. Verify the URL is correct.
Performance Tips
- Implement caching: Cache frequently accessed pages to reduce API calls
- Use parallel processing: Fetch multiple pages concurrently (with rate limiting)
- Process content efficiently: Extract only needed information from HTML
- Set timeouts: Implement reasonable timeouts for page fetching
- Monitor token usage: Track usage to optimize costs
- Batch operations: Group multiple URL fetches when possible
Security Considerations
- Validate all URLs before processing
- Sanitize extracted HTML content before displaying
- Implement rate limiting to prevent abuse
- Never expose SDK credentials in client-side code
- Be respectful of robots.txt and website terms of service
- Handle user data according to privacy regulations
- Implement proper error handling for failed requests
Remember
- Always use z-ai-web-dev-sdk in backend code only
- The SDK is already installed - import as shown in examples
- Implement proper error handling for robust applications
- Use caching to improve performance and reduce costs
- Respect website terms of service and rate limits
- Process HTML content carefully to extract meaningful data
- Monitor token usage for cost optimization


