web-scraper
SolidВежливый скраппинг HTML-страниц в Markdown/JSON. Скрипт scrape.py читает страницу по URL, применяет простой CSS-селектор (tag, tag#id, tag.class) для выбора строк/секций, извлекает текст, ссылки и таблицы и отдаёт результат в Markdown или JSON. Легальные guardrails встроены в код: проверка robots.txt, честный User-Agent, задержка между запросами (по умолчанию 1.0 с), лимит размера страницы 10 МБ. Триггеры: 'web scraping', 'скраппинг', 'скачать данные с сайта', 'парсинг сайта', 'парсер html', 'извлечь данные', 'scrape', 'scraping'.
Install
Quality Score: 82/100
Skill Content
Details
- Author
- bestdeejay-design
- Repository
- bestdeejay-design/agent-skills
- Created
- 1 months ago
- Last Updated
- 2 weeks ago
- Language
- Python
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
web-scraper
Generate browser console scripts to scrape paginated websites. Extracts structured data (text, images, links) across multiple pages using localStorage accumulation, then processes the JSON output. Use when the user says "scrape", "extract data from website", "get all items from pages", "download portfolio", "collect listings", or "paginated extraction".
web-scrape
Fetch and parse web content with ethical scraping practices, rate limiting, and structured extraction
scrapling
Guides safe, version-aware use of Scrapling for HTML extraction, adaptive selectors, static or browser-backed fetching, spiders, proxy rotation, robots.txt-aware crawling, RAG markdown generation, and MCP integration. Use when choosing a Scrapling fetcher/session, building a crawl, generating LLM-ready markdown, repairing selectors after site changes, migrating from BeautifulSoup/Scrapy, or validating Scrapling API usage against the repository's pinned upstream version.