一套覆盖 14 章的完整学习体系:从环境搭建与 HTTP 原理,到请求库、解析库、并发、动态渲染、反爬对抗、Scrapy 框架、接口逆向,再到工程化与多个完整实战项目。图文结合,代码全部可直接运行。
爬虫(Web Crawler / Spider)的本质只有一句话:用代码模拟浏览器发出 HTTP 请求,把服务器返回的数据结构化地保存下来。搜索引擎、价格监控、舆情分析、数据日报,背后都是爬虫。它不神秘,本身也不违法——关键在于爬什么、怎么爬、用来做什么。
这四个环节循环往复。一个成熟的爬虫,就是在每个环节上不断加固:请求要能突破拦截,解析要能从脏 HTML 里精准提取,存储要能扛住数据量和中断。
| 类型 | 目标 | 典型场景 | 技术重点 |
|---|---|---|---|
| 通用爬虫 | 全网抓取 | 搜索引擎 | 调度、去重、分布式 |
| 聚焦爬虫 | 特定站点/主题 | 竞品监控、数据日报 | 解析规则、字段抽取 |
| 增量爬虫 | 只抓新增/变更 | 新闻订阅、价格变动 | 去重、断点、时间戳 |
| 深层爬虫 | 登录后/动态内容 | 后台数据、SPA 应用 | 会话、渲染、逆向 |
本教程聚焦聚焦爬虫与深层爬虫——这是个人和中小团队最常用的两类,也是投入产出比最高的方向。
合法场景依然很宽广:公开数据采集、学术研究、市场分析、自用工具、SEO 监测。把力气花在技术上、把边界记在心里,才是这一行能长期做下去的前提。
工欲善其事,必先利其器。本章搭建一套可长期使用的爬虫开发环境。
推荐使用 Python 3.12+(3.11 仍受支持,也够用)。包管理用 venv、conda 或 uv 均可,关键是把爬虫依赖装进隔离环境,避免污染系统 Python。uv 是近两年最流行的新选择,装库快很多。
externally-managed-environment?这不是故障——是 PEP 668 在提醒你「别直接往系统环境装包」,按下面命令先建虚拟环境即可。确需强装才加 --break-system-packages,不推荐。# 查看 Python 版本(建议 3.11+)
python --version
# 创建隔离环境(conda 示例)
conda create -n crawl python=3.11 -y
conda activate crawl
# 一站式安装(按需求选择)
pip install scrapling # 一体化抓取框架(本教程主力)
pip install requests # 经典请求库
pip install httpx # 现代请求库(HTTP/2 + 异步)
pip install beautifulsoup4 lxml # 解析库
pip install parsel # Scrapy 同款选择器库
pip install playwright # 浏览器渲染
pip install scrapy # 爬虫框架(进阶)
python -m playwright install chromium
python -m patchright install chromium
scrapling 的 DynamicFetcher 走 playwright,StealthyFetcher 走 patchright,两套浏览器是分开的,建议都装上。VS Code / PyCharm。装 Python 扩展,开启类型提示,写爬虫效率高很多。
Chrome/Edge。F12 打开开发者工具:Elements 看结构、Network 看请求、Console 调试 JS。
进阶用 mitmproxy / Charles / Fiddler 抓 App 或复杂网页请求。
DB Browser for SQLite(看 .db 文件)、Navicat(看 MySQL)。
跑通下面这段最小示例,环境就算就绪:
from scrapling.fetchers import Fetcher
page = Fetcher.get('https://example.com')
print('状态码:', page.status) # 200
print('标题:', page.css('h1::text').get()) # Example Domain
200 和 Example Domain。如果报 AttributeError: ... no attribute 'fetch',说明你用了 Fetcher.fetch()——静态抓取请用 .get()(详见第 5 章)。爬虫工程师 90% 的时间不是在写代码,而是在理解"数据是怎么从服务器到页面的"。HTTP 是这一切的底层,必须吃透。
一个 URL 拆解后,每一段对爬虫都有意义:
https://demo-site.com:443/browse?finished=1&sort=release_date#page2
└─┬─┘ └──────┬──────┘└┬┘└─┬──┘└──────────┬──────────┘└─┬──┘
协议 域名 端口 路径 查询参数 锚点
| 部分 | 含义 | 爬虫视角 |
|---|---|---|
| 协议 | http / https | https 需注意证书与 TLS 指纹 |
| 域名 | 目标主机 | 子域名常有独立数据接口 |
| 路径 | 资源位置 | /anime/{id} 这类可遍历 |
| 查询参数 | ?key=val&key=val | 分页、筛选、排序都藏在参数里,改参数=改数据 |
| 锚点 | #xxx | 只在前端生效,服务器收不到,抓取时可忽略 |
?page=2、?sort=price、?category=tech 这类参数,直接改它们就能批量抓不同数据。| 方法 | 用途 | 是否带请求体 | 爬虫常见度 |
|---|---|---|---|
| GET | 获取资源 | 否(参数在 URL) | ★★★★★ |
| POST | 提交数据 | 是 | ★★★(登录、表单、接口) |
| PUT / PATCH | 更新资源 | 是 | ★(多见于 API) |
| DELETE | 删除资源 | 视情况 | ★(少见) |
爬虫主要用 GET。遇到登录、搜索、复杂查询时才会用 POST(数据放在请求体里,常见 JSON 或表单格式)。
请求头是反爬攻防的第一战场。关键字段逐个说:
| 请求头 | 作用 | 反爬相关性 |
|---|---|---|
User-Agent | 声明客户端身份 | 最高——裸 requests 的 UA 一眼被识破 |
Referer | 从哪个页面跳转来 | 高——防盗链、图片、视频常校验 |
Cookie | 登录态/会话标识 | 高——登录后数据全靠它 |
Accept | 期望的响应格式 | 中——application/json vs text/html |
Accept-Language | 语言偏好 | 低——部分站点据此返回中英文 |
Content-Type | 请求体格式(POST) | 中——POST 时后端会校验 |
X-Requested-With | 标识 AJAX 请求 | 中——XMLHttpRequest 常见于接口 |
import requests
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36',
'Referer': 'https://www.google.com/',
'Accept': 'text/html,application/xhtml+xml',
'Accept-Language': 'zh-CN,zh;q=0.9',
}
r = requests.get('https://example.com', headers=headers, timeout=10)
print(r.status_code)
| 状态码 | 含义 | 爬虫应对策略 |
|---|---|---|
| 200 | 成功 | 正常解析响应体 |
| 301 / 302 | 永久 / 临时重定向 | requests 默认跟随;注意最终 URL |
| 304 | 未修改(缓存) | 配合 If-Modified-Since 做增量 |
| 400 | 请求参数错误 | 检查 URL 参数、请求体格式 |
| 401 | 未授权 | 需要登录/Token |
| 403 | 禁止访问 | 被反爬拦了:查 UA、加 Referer、换代理 |
| 404 | 资源不存在 | 链接失效,记录跳过 |
| 429 | 请求过于频繁 | 立即限速:加长随机延时、退避重试 |
| 500 / 502 / 503 | 服务器错误 | 稍后重试;503 常是 WAF 拦截 |
HTTP 本身无状态,服务器靠 Cookie 记住"你是谁"。理解这个机制,才能处理登录态。
对应到爬虫:登录一次拿到 Cookie 后,用会话对象保持它,后续请求就都是"已登录"状态。requests 的 Session、scrapling 的会话都自动处理了 Cookie 的存取。
现在的网站基本都是 HTTPS。加密握手时,客户端会暴露一组特征(支持的加密套件、扩展、顺序等),即 TLS 指纹。普通 Python 请求的 TLS 指纹和真实浏览器完全不同——这就是为什么有些站点 requests 一请求就 403,而浏览器正常。
curl_cffi 或 scrapling 这类能模拟浏览器 TLS 指纹的库。这是 2024 年后反爬对抗的关键技术点,也是本教程推荐 scrapling 的重要原因之一。import requests
r = requests.get('https://httpbin.org/headers', headers={'User-Agent': 'MyBot/1.0'})
print('请求头回显:')
print(r.json()) # 服务器收到的所有请求头
print('响应状态:', r.status_code)
print('响应头:', dict(r.headers))
print('编码:', r.encoding)
httpbin.org 是一个专门用于测试 HTTP 的公开服务,会把你的请求头原样返回——用它练习观察请求头是绝佳方式。
拿到 HTML 后,怎么精准定位到想要的那一小块数据?靠选择器。这是爬虫的"手术刀"。
HTML 本质是嵌套的标签树。浏览器 F12 → Elements 面板看到的就是这棵树的实时渲染。爬虫要做的就是:从根节点出发,用选择器定位到目标节点,取出它的文本或属性。
| 需求 | 写法 | 说明 |
|---|---|---|
| 按标签名 | a | 所有 a 标签 |
| 按 class | .title | class 含 title |
| 按 id | #main | id 为 main(唯一) |
| 按属性 | a[href]、img[data-src] | 有某属性即可 |
| 属性值 | a[href="/login"] | 属性等于某值 |
| 属性开头 | a[href^="/anime/"] | href 以 /anime/ 开头 |
| 属性结尾 | img[src$=".jpg"] | src 以 .jpg 结尾 |
| 属性包含 | div[class*="list"] | class 包含 list |
| 后代 | div p | div 内所有 p(任意层级) |
| 直接子级 | div > p | div 的直接子 p |
| 第 n 个 | li:nth-child(2) | 父级下第 2 个 li |
| 取文本 | h1::text | 伪选择器,取文本 |
| 取属性 | img::attr(src) | 伪选择器,取属性值 |
::text、::attr(x) 是 parsel / scrapling 的扩展伪选择器,不是标准 CSS。在 BeautifulSoup 里要用 .get_text() 和 ['src'] 取。XPath 是另一种定位语言,在处理按文本内容匹配、向上查找父节点、复杂条件时比 CSS 更强大。
| 需求 | XPath | 说明 |
|---|---|---|
| 所有 a | //a | 双斜杠=任意层级 |
| 按 class | //*[@class="title"] | @取属性 |
| 按 id | //*[@id="main"] | |
| 按文本 | //a[text()="登录"] | CSS 做不到 |
| 文本包含 | //a[contains(text(),"下载")] | 模糊匹配 |
| 取属性 | //img/@src | @前缀 |
| 取文本 | //h1/text() | |
| 父节点 | //a/.. | 向上回溯 |
| 兄弟节点 | //h2/following-sibling::p | 同级之后 |
| 多条件 | //div[@class="a" and @data-x="1"] | and/or 组合 |
短、快、可读性好。80% 的场景 CSS 就够了。
按文本内容定位(如"找文字含'价格'的标签")、需要向上找父节点、复杂多条件。
当数据混在 JS 代码里、或 HTML 结构极不规则时,正则是最后的兜底手段。
import re
html = '价格:¥1299 元,库存:42 件'
price = re.search(r'¥(\d+)', html).group(1) # 1299
stock = re.search(r'库存:(\d+)', html).group(1) # 42
# 常用正则片段
# \d+ 数字 \w+ 字母数字 .* 任意 \s+ 空白 [^"]+ 非引号
F12 → 左上角箭头图标 → 点击页面上的目标,自动跳到对应 HTML 节点。
Elements 面板按 Ctrl+F,输入选择器表达式,实时高亮匹配结果,确认选对再写进代码。
右键节点 → Copy → Copy selector / Copy XPath,直接得到可用表达式(但常偏长,需精简)。
右键 → View page source 看源码。目标数据在源码里=静态可直抓;不在=JS 渲染,转第 9 章。
请求库负责"发出请求、拿回响应"。选对库、用对方法,能省掉一大半反爬的麻烦。
| 库 | 定位 | 优势 | 短板 | 何时用 |
|---|---|---|---|---|
| urllib | 标准库 | 无需安装 | API 繁琐 | 基本不用,了解即可 |
| requests | 经典同步 | 简单直观、生态最好 | TLS 指纹易被识别 | 学习原理、简单静态站 |
| httpx | 现代同步/异步 | HTTP/2、异步、API 像 requests | 同样有指纹问题 | 要 HTTP/2 或异步时 |
| curl_cffi | 指纹伪装 | 完美模拟浏览器 TLS 指纹 | 稍重 | 反爬严格的静态站 |
| scrapling | 一体化框架 | 请求+解析+动态+反爬一体 | 较新 | 本教程主力 |
import requests
# GET 带参数(自动拼到 URL)
r = requests.get('https://httpbin.org/get',
params={'page': 2, 'sort': 'date'},
headers={'User-Agent': 'Mozilla/5.0'},
timeout=10)
print(r.url) # https://httpbin.org/get?page=2&sort=date
print(r.status_code) # 200
print(r.text[:100]) # 文本内容
print(r.json()) # 若响应是 JSON 直接转字典
# POST 提交表单 / JSON
requests.post(url, data={'user': 'a'}) # 表单格式
requests.post(url, json={'user': 'a'}) # JSON 格式
用 Session 对象,它会自动在多次请求间保持 Cookie:
s = requests.Session()
# 第一步:登录(Cookie 自动存入 Session)
s.post('https://site.com/login', data={'user': 'a', 'pwd': 'b'})
# 第二步:之后的请求都带着登录态
r = s.get('https://site.com/profile') # 已登录,能看到私密数据
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
s = requests.Session()
retries = Retry(total=3, backoff_factor=1,
status_forcelist=[429, 500, 502, 503])
s.mount('https://', HTTPAdapter(max_retries=retries))
r = s.get(url, timeout=10) # 超 10 秒自动放弃,遇 5xx/429 自动重试 3 次
proxies = {
'http': 'http://127.0.0.1:7890',
'https': 'http://127.0.0.1:7890',
}
r = requests.get(url, proxies=proxies, timeout=10)
127.0.0.1),requests / curl_cffi 会自动走系统代理,可能导致部分站点连不上或走弯路。遇到诡异的 over proxy 127.0.0.1 连接错误,先检查系统代理,必要时显式传 proxies={} 禁用,或换用浏览器模式。scrapling 把请求和解析合在一起,且自动处理 TLS 指纹。三种 Fetcher 对应三种场景:
| Fetcher | 方法 | 底层 | 适用 |
|---|---|---|---|
Fetcher | .get() / .post() | curl_cffi | 静态页面,最快 |
DynamicFetcher | .fetch() | Playwright | JS 渲染页 |
StealthyFetcher | .fetch() | patchright | 强反爬(Cloudflare) |
from scrapling.fetchers import Fetcher, DynamicFetcher, StealthyFetcher
# ① 静态:用 get,不是 fetch!
page = Fetcher.get('https://example.com')
# ② 动态渲染
page = DynamicFetcher.fetch('https://spa-site.com', headless=True)
# ③ 强反爬
page = StealthyFetcher.fetch('https://protected.com', headless=True,
network_idle=True)
Fetcher(静态)只有 .get()/.post(),没有 .fetch();两个浏览器 Fetcher 才有 .fetch()。写错就报 AttributeError: type object 'Fetcher' has no attribute 'fetch'。记住:静态用 get,浏览器用 fetch。from curl_cffi import requests as creq
# impersonate 指定模拟哪个浏览器的 TLS 指纹
r = creq.get('https://example.com', impersonate='chrome')
print(r.status_code)
拿到 HTML 后,用解析库把数据"抠"出来。主流四个:BeautifulSoup、lxml、parsel、正则。
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'lxml')
title = soup.find('h1').get_text() # 按标签找
items = soup.select('.item a') # CSS 选择器
for a in items:
print(a.get_text(), a['href']) # 文本 + 属性
# 按文本找(BeautifulSoup 也支持)
tag = soup.find('a', string='登录')
from lxml import etree
tree = etree.HTML(html)
titles = tree.xpath('//h1/text()') # XPath
links = tree.xpath('//a[@class="item"]/@href') # 取属性
from parsel import Selector
sel = Selector(html)
sel.css('h1::text').get() # CSS 取文本
sel.xpath('//a/@href').getall() # XPath 取属性
sel.css('a').re(r'href="(.*?)"') # 甚至能套正则
import json
data = r.json() # requests 直接转
# 或
data = json.loads(text) # 字符串转字典
title = data['data']['list'][0]['title'] # 按层级取值
# 嵌套深时用 jmespath 更优雅
import jmespath
titles = jmespath.search('data.list[*].title', data)
| 方式 | 上手难度 | 速度 | 功能 | 推荐场景 |
|---|---|---|---|---|
| BeautifulSoup | ★ 最简单 | 中 | CSS + 部分文本查找 | 入门、快速验证 |
| lxml | ★★ | 最快 | 纯 XPath | 追求性能 |
| parsel | ★★ | 快 | CSS+XPath+正则 | Scrapy 生态 |
| 正则 | ★★★ | 快 | 任意文本 | JS 抠数据、兜底 |
from scrapling.fetchers import Fetcher
page = Fetcher.get('https://demo-site.com/')
# 一次抓取多个字段,组合成结构化数据
for card in page.css('a[href^="/anime/"]'):
href = card.attrib.get('href')
name = card.css('img::attr(alt)').get()
if name:
print(name, href)
数据抓下来要落地。按数据量和查询需求选存储方式。
import csv, json
# CSV:表格数据,Excel 直接打开(utf-8-sig 防乱码)
with open('data.csv', 'w', newline='', encoding='utf-8-sig') as f:
w = csv.DictWriter(f, fieldnames=['title', 'price'])
w.writeheader()
w.writerows(rows)
# JSON:结构化、嵌套数据
json.dump(data, open('data.json', 'w', encoding='utf-8'),
ensure_ascii=False, indent=2)
# JSONL:一行一条,适合大数据量逐条追加
with open('data.jsonl', 'a', encoding='utf-8') as f:
f.write(json.dumps(item, ensure_ascii=False) + '\n')
零配置、单文件、支持 SQL,个人爬虫首选。
import sqlite3
conn = sqlite3.connect('data.db')
conn.execute('''CREATE TABLE IF NOT EXISTS anime(
id INTEGER PRIMARY KEY AUTOINCREMENT,
title TEXT, url TEXT UNIQUE, meta TEXT)''')
# 插入(UNIQUE 约束 + INSERT OR IGNORE 自动去重)
conn.execute('INSERT OR IGNORE INTO anime(title,url,meta) VALUES (?,?,?)',
(title, url, meta))
conn.commit()
# 查询
for row in conn.execute('SELECT title, url FROM anime LIMIT 5'):
print(row)
| 存储 | 适用规模 | 特点 | 何时用 |
|---|---|---|---|
| CSV / JSON | < 10 万条 | 最简单、可读 | 一次性小任务 |
| SQLite | < 百万条 | 零配置、单文件 | 个人项目首选 |
| MySQL | 百万+ | 关系型、多表关联 | 结构化、需复杂查询 |
| MongoDB | 百万+ | 文档型、灵活 | 字段不固定、JSON 类数据 |
| Redis | 内存级 | 极快、支持队列/集合 | 去重集合、任务队列 |
import redis
r = redis.Redis(host='localhost', port=6379, db=0)
# 集合去重:sadd 返回 1=新增,0=已存在
if r.sadd('crawled', url):
# 新 URL,加入待抓队列
r.lpush('todo', url)
# 从队列取任务
task = r.rpop('todo')
串行抓 1 万条可能要几小时,并发能压到几分钟。但并发是把双刃剑——提速的同时也更容易被封。
网络请求是 I/O 密集型,多线程提速明显、代码改动最小:
from concurrent.futures import ThreadPoolExecutor
from scrapling.fetchers import Fetcher
urls = [f'https://site.com/anime/{i}' for i in range(100, 130)]
def fetch(url):
try:
return url, Fetcher.get(url).css('title::text').get()
except Exception as e:
return url, f'ERR: {e}'
with ThreadPoolExecutor(max_workers=5) as ex: # 5 并发,别贪多
results = list(ex.map(fetch, urls))
for url, title in results:
print(title)
import asyncio, aiohttp
async def fetch(session, url):
async with session.get(url) as resp:
return url, await resp.text()
async def main():
async with aiohttp.ClientSession() as s:
tasks = [fetch(s, u) for u in urls]
return await asyncio.gather(*tasks)
results = asyncio.run(main())
# 限制同时只有 5 个请求在飞,避免瞬间打满
sem = asyncio.Semaphore(5)
async def fetch(session, url):
async with sem: # 拿到令牌才执行
async with session.get(url) as resp:
return await resp.text()
| 场景 | 推荐 | 理由 |
|---|---|---|
| 入门 / 小任务(< 1 千) | 串行 | 简单、不易被封、好调试 |
| 中等(1 千–10 万) | 多线程 | 提速明显、代码改动小 |
| 大规模(10 万+) | 协程 / Scrapy | 性能与可维护性 |
当数据由 JavaScript 动态生成(源码里看不到、Network 里有 XHR),就进入动态渲染的领域。
很多"动态站"其实是前端渲染 + 后端 JSON 接口。直接请求 JSON 接口,比开浏览器快一个数量级、省资源、还稳定。
右键 → View page source,搜目标数据。源码里有=静态直抓;没有=动态。
F12 → Network → 刷新 → 筛选 Fetch/XHR,逐个看响应,找到返回目标数据的那个请求。
右键该请求 → Copy as cURL → 转 Python。直接请求这个接口拿 JSON,绕过渲染。
才用浏览器渲染(Playwright / DynamicFetcher)。
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto('https://spa-site.com')
# 等待关键元素加载完成(动态页必须)
page.wait_for_selector('.product-list', timeout=10000)
# 滚动加载更多(无限滚动页面)
page.mouse.wheel(0, 3000)
page.wait_for_timeout(2000)
html = page.content() # 拿渲染后的完整 HTML
page.screenshot(path='shot.png') # 截图调试用
browser.close()
from scrapling.fetchers import DynamicFetcher, StealthyFetcher
# Dynamic:JS 渲染后拿页面
page = DynamicFetcher.fetch(
'https://spa-site.com',
headless=True,
network_idle=True, # 等网络空闲,确保 JS 加载完
wait_selector='.product', # 等特定元素出现
)
print(page.css('h1::text').get())
# Stealthy:强反爬(Cloudflare 盾等),真实浏览器指纹
page = StealthyFetcher.fetch(
'https://protected.com',
headless=True,
solve_cloudflare=True, # 自动过 Turnstile 盾
)
solve_cloudflare 这类能力只应用于你自有或已获授权的站点(比如测试自己站点的防护配置)。把它用在他人站点上绕过安全防护,很可能踩中第 1.3 节讲的法律红线——技术上能做 ≠ 做了没事。学习中遇到强防护站点,正确做法是换数据源或用官方 API,而不是硬破。scrapling 的 capture_xhr 能拦截页面发出的接口请求——既不用解析渲染后 HTML,也不用自己逆向接口:
page = DynamicFetcher.fetch(
'https://spa-site.com/products',
capture_xhr='/api/products', # 拦截含此关键字的请求
)
for req in page.captured_xhr:
data = req.json() # 直接拿到接口返回的 JSON
print(data)
capture_xhr 是"路径关键字匹配拦截",不是"强制抓取"。只有页面真实发出了含该关键字的请求才会被拦下。匹配不到就返回空(结果里不会出现该字段)——纯 SSR 站点(无 XHR)传这个参数没用。Selenium 是浏览器自动化的老前辈,语法啰嗦、速度慢,现在基本被 Playwright 取代。除非维护老项目,否则直接学 Playwright。
反爬是猫鼠游戏。理解对方的检测手段,才知道怎么应对。本章按"从轻到重"梳理完整对抗体系。
| 层级 | 反爬手段 | 你的应对 |
|---|---|---|
| 1 | 检查 User-Agent | 设置真实浏览器 UA |
| 2 | 校验 Referer(防盗链) | 补全 Referer,模拟站内跳转 |
| 3 | 频率限制 / IP 封禁 | 随机延时、指数退避、限速 |
| 4 | 需要 Cookie / 登录 | Session 保持登录态 |
| 5 | TLS / 浏览器指纹检测 | curl_cffi / StealthyFetcher 模拟指纹 |
| 6 | 滑块、验证码、人机验证 | 成本极高,建议换数据源 |
import time, random
for url in urls:
fetch(url)
time.sleep(random.uniform(1, 3)) # 随机 1-3 秒,模拟人类
proxy_list = [
'http://1.2.3.4:8080',
'http://5.6.7.8:8080',
]
proxy = random.choice(proxy_list)
r = requests.get(url, proxies={'http': proxy, 'https': proxy})
自建代理池需维护可用性检测(定期测速、剔除失效 IP),也可用付费代理服务。小规模抓取用本地代理(如 Clash)即可。
把抓包结果、混淆 JS、异常响应喂给大模型,让它帮你分析"参数怎么生成的""哪里被拦了"。这是近两年爬虫工作流的重大变化——逆向分析的效率被 AI 大幅拉高。
当爬虫需要规模化、规范化、长期维护时,用 Scrapy 框架。它把"请求调度、解析、存储、去重、中间件"都标准化了。
# 命令行创建项目
scrapy startproject myspider
cd myspider
scrapy genspider anime demo-site.com # 生成一个 spider
scrapy crawl anime # 运行
import scrapy
class AnimeSpider(scrapy.Spider):
name = 'anime'
start_urls = ['https://demo-site.com/']
def parse(self, response):
# 解析列表页,提取详情链接并跟进
for href in response.css('a[href^="/anime/"]::attr(href)').getall():
yield response.follow(href, self.parse_detail)
def parse_detail(self, response):
yield {
'title': response.css('h1::text').get(),
'url': response.url,
'meta': response.css('meta[name="description"]::attr(content)').get(),
}
# pipelines.py · 存入 SQLite
import sqlite3
class SQLitePipeline:
def open_spider(self, spider):
self.conn = sqlite3.connect('anime.db')
self.conn.execute('CREATE TABLE IF NOT EXISTS anime(title TEXT,url TEXT,meta TEXT)')
def process_item(self, item, spider):
self.conn.execute('INSERT INTO anime VALUES (?,?,?)',
(item['title'], item['url'], item['meta']))
self.conn.commit()
return item
USER_AGENT = 'Mozilla/5.0 ...'
DOWNLOAD_DELAY = 1 # 请求间隔(秒),礼貌限速
CONCURRENT_REQUESTS = 8 # 全局并发数
ROBOTSTXT_OBEY = True # 遵守 robots.txt
ITEM_PIPELINES = {'myspider.pipelines.SQLitePipeline': 300}
很多高价值数据藏在接口里。学会抓包分析接口,是爬虫进阶的分水岭。
| 工具 | 特点 | 适用 |
|---|---|---|
| 浏览器 DevTools | 内置、零安装 | 网页接口分析(首选) |
| mitmproxy | 开源、可脚本化 | App / 复杂网页抓包 |
| Charles / Fiddler | GUI 友好 | 桌面端抓包 |
Network → Fetch/XHR,逐个看响应,找到返回目标数据的请求。
看 Query 参数和请求体:哪些是固定的(如页码),哪些是动态生成的(如 sign、token、timestamp)。
固定参数直接复制;动态签名看 JS 怎么生成(可用 LLM 辅助读混淆代码)。
Copy as cURL → 转 Python → 替换动态参数 → 循环调用。
# 常见签名 = md5(参数拼接 + 密钥 + 时间戳)
import hashlib, time
def make_sign(params, secret):
s = ''.join(f'{k}={v}' for k, v in sorted(params.items()))
s += secret + str(int(time.time()))
return hashlib.md5(s.encode()).hexdigest()
params = {'page': 1, 'size': 20}
params['sign'] = make_sign(params, '网站JS里抠出来的密钥')
r = requests.get('https://api.site.com/list', params=params)
sign、md5、encrypt 等关键词定位生成逻辑,或直接把 JS 片段交给 LLM 分析。App 的数据接口往往比网页更干净。用 mitmproxy 抓手机流量(手机设代理指向电脑),分析 App 的接口。注意:现代 App 多有证书绑定(SSL Pinning),需额外手段绕过,门槛较高,学习阶段建议先用网页版练手。
抓几十条不需要工程化;抓几万条还要长期稳定跑,就需要下面这些"基础设施"。
核心思想:已抓的 URL 存进数据库,启动时先读出来跳过。配合异常捕获,就能跑几天几夜不中断、崩了也不白干。
import sqlite3, time, random
conn = sqlite3.connect('crawl.db')
conn.execute('CREATE TABLE IF NOT EXISTS done(url TEXT PRIMARY KEY)')
conn.execute('CREATE TABLE IF NOT EXISTS anime(url TEXT, title TEXT)')
done = {r[0] for r in conn.execute('SELECT url FROM done')}
for url in urls:
if url in done: continue # 断点跳过
try:
title = Fetcher.get(url).css('title::text').get()
conn.execute('INSERT INTO anime VALUES (?,?)', (url, title))
conn.execute('INSERT INTO done VALUES (?)', (url,))
conn.commit() # 每成功一条立即落库
except Exception as e:
print('失败跳过:', url, e)
time.sleep(random.uniform(0.5, 1.5))
只抓新增/变更数据,不重复抓全量。常用两种策略:
import logging
logging.basicConfig(
level=logging.INFO,
format='[%(asctime)s] %(levelname)s: %(message)s',
handlers=[logging.FileHandler('crawl.log', encoding='utf-8'),
logging.StreamHandler()]
)
log = logging.getLogger('crawl')
log.info('抓取成功 %s', url)
log.warning('第 %d 次重试', attempt)
log.error('失败 %s: %s', url, e)
| 方式 | 命令/配置 | 适用 |
|---|---|---|
| Windows 任务计划 | taskschd.msc 新建任务 | Windows 本机定时 |
| Linux cron | 0 8 * * * python crawl.py | 服务器定时 |
| 云函数 | 腾讯云函数等定时触发 | 免维护、按需付费 |
| APScheduler | 代码内定时 | 需常驻进程 |
# APScheduler 示例(常驻进程内定时)
from apscheduler.schedulers.blocking import BlockingScheduler
sched = BlockingScheduler()
sched.add_job(crawl_task, 'cron', hour=8) # 每天 8 点跑
sched.start()
四个由浅入深的完整项目,覆盖本教程所有核心技能。
目标:稀饭动漫 Next,抓首页番剧链接 → 详情页 → 存 SQLite。用到:静态抓取、列表解析、翻页、断点续爬、限速。
# crawl_anime.py
import sqlite3, time, random
from scrapling.fetchers import Fetcher
BASE = 'https://demo-site.com'
conn = sqlite3.connect('anime.db')
conn.execute('CREATE TABLE IF NOT EXISTS anime(id INTEGER PRIMARY KEY,title TEXT,url TEXT,meta TEXT)')
conn.execute('CREATE TABLE IF NOT EXISTS done(url TEXT PRIMARY KEY)')
def links_of(page):
return set(page.css('a[href^="/anime/"]::attr(href)').getall())
def save(href):
if conn.execute('SELECT 1 FROM done WHERE url=?',(href,)).fetchone(): return
p = Fetcher.get(BASE+href)
conn.execute('INSERT INTO anime(title,url,meta) VALUES (?,?,?)',
(p.css('h1::text').get(), href,
p.css('meta[name="description"]::attr(content)').get()))
conn.execute('INSERT INTO done VALUES (?)',(href,)); conn.commit()
links = set()
for i in range(1,4):
links |= links_of(Fetcher.get(f'{BASE}/browse?page={i}'))
time.sleep(random.uniform(1,2))
for h in sorted(links):
try: save(h)
except Exception as e: print('跳过', h, e)
print('入库:', conn.execute('SELECT COUNT(*) FROM anime').fetchone()[0])
目标:定时抓商品页价格,入库对比涨跌。用到:多线程、SQLite、定时调度、增量。
import sqlite3, time
from concurrent.futures import ThreadPoolExecutor
from scrapling.fetchers import Fetcher
URLS = ['https://shop.com/item/1', 'https://shop.com/item/2']
conn = sqlite3.connect('price.db')
conn.execute('CREATE TABLE IF NOT EXISTS price(ts REAL,item TEXT,price REAL)')
def grab(u):
p = Fetcher.get(u)
price = p.css('.price::text').get()
conn.execute('INSERT INTO price VALUES (?,?,?)',(time.time(),u,price)); conn.commit()
return u, price
with ThreadPoolExecutor(3) as ex:
for u, price in ex.map(grab, URLS): print(u, price)
目标:抓页面所有图片地址并下载。用到:属性提取、文件流写入、并发下载。
import os, requests
from concurrent.futures import ThreadPoolExecutor
from scrapling.fetchers import Fetcher
os.makedirs('imgs', exist_ok=True)
page = Fetcher.get('https://site.com')
imgs = page.css('img::attr(src)').getall()
def dl(src):
if not src.startswith('http'): return
name = src.split('/')[-1]
r = requests.get(src, timeout=15)
open(f'imgs/{name}', 'wb').write(r.content)
with ThreadPoolExecutor(8) as ex:
ex.map(dl, imgs)
目标:抓一个前端渲染的列表页,直接截获其接口 JSON。用到:DynamicFetcher、capture_xhr。
from scrapling.fetchers import DynamicFetcher
page = DynamicFetcher.fetch('https://spa-site.com/list',
headless=True, capture_xhr='/api/list')
for req in page.captured_xhr:
for item in req.json().get('data', []):
print(item)
| 项目 | 请求 | 解析 | 存储 | 并发 | 进阶点 |
|---|---|---|---|---|---|
| 动漫目录 | 静态 | CSS | SQLite | — | 断点续爬 |
| 价格监控 | 静态 | CSS | SQLite | 多线程 | 定时增量 |
| 图片下载 | 静态 | 属性 | 文件 | 多线程 | 文件流 |
| SPA 数据 | 渲染 | JSON | — | — | XHR 截获 |
| 问题 | 原因 | 解决 |
|---|---|---|
AttributeError: no attribute 'fetch' | 静态 Fetcher 用了 fetch | 改用 .get() |
over proxy 127.0.0.1 连接失败 | 系统代理拦截 | 显式 proxies={} 或换浏览器模式 |
| 抓到的数据是空的 | JS 渲染,源码无数据 | 转第 9 章动态渲染 |
| 突然 403 / 429 | 触发频率限制 | 加长延时、降并发、换 IP |
| 中文乱码 | 编码问题 | 设 r.encoding 或用 utf-8-sig 存 CSV |
| capture_xhr 为空 | 站点无 XHR(纯 SSR) | 该站静态抓即可 |
《Python 3 网络爬虫开发实战(第2版)》——覆盖全流程,中文经典。
Playwright Docs、Scrapy 教程、parsel、curl_cffi、scrapling README。
httpbin.org(测请求)、公开 API 列表、GitHub 高星爬虫项目。
curlconverter(cURL 转代码)、DB Browser for SQLite、mitmproxy。
scraplingrequestshttpxcurl_cffi parselBeautifulSouplxmljmespath playwrightscrapyaiohttpsqlite3 redisapschedulermitmproxy