Python后端爬虫专题26:测试不是“请求一下看看”——Fixture、RESPX与端到端旅程
上一篇练习完整答案
target.test.evil.example的 hostname 不在精确 allowlist,入口即拒绝;target.test@evil.example中 target.test 是 userinfo,真实 hostname 为 evil.example,同样在 URL 策略拒绝;127.0.0.1 是字面回环地址,即使错误加入 host allowlist,IP 分类仍拒绝;允许域 302 到内网时,初始请求可以发送,但解析 Location 后必须重新做 host、DNS 地址检查,在第二跳前停止。
超大响应的完整行为已由tests/test_http_client.py::test_fetcher_stops_reading_when_response_exceeds_budget覆盖:MockTransport 返回超过预算的 body,HttpFetcher 流式计数并抛 ResponseTooLarge,调用者没有机会进入 parse_detail。三层出站设计:应用仅允许精确 TargetLab;Compose 网络只让 Worker 访问 targetlab/postgres/redis/minio 所需端口,不暴露任意出口;云/宿主防火墙拒绝元数据和内网段。allowed_private_hosts=targetlab只因课程服务位于 Docker 私网,不能复制到用户可控 host。
一条失败测试应告诉你哪一层坏了
解析器测试使用固定 HTML Fixture,不走网络;HTTP 客户端测试用 MockTransport/RESPX 控制状态码、头和重定向;Repository 用临时 SQLite;API 用 ASGITransport;端到端旅程把真实 TargetLab、HttpFetcher、Crawler、快照与数据库连起来。层次越低越快、定位越准,层次越高越能发现接口拼接错误。
只有端到端测试,失败时很难知道是 CSS、重试还是数据库;只有单元测试,各模块都绿却可能没有正确接线。课程两者都保留。
Fixture应该像真实输入,不像实现抄写
jobs-page-1.html包含两个详情、跟踪参数和下一页;详情 Fixture 有标题、公司、城市、薪资、日期、多段描述、技能;坏详情专门缺 title。测试期望使用手工确定的规范 URL 与字段,不调用 normalize_url 生成 expected,否则实现和期望可能一起错。
页面改版时,不要立刻改掉旧 Fixture。先把线上失败快照脱敏后加入回归样本,写出旧/新结构的产品决策:兼容两种还是明确切换。Fixture 文件名可含来源、场景和日期,避免test1.html。
RESPX/MockTransport负责协议场景
真实公网不能稳定地产生“第一次503、第二次200”“429带Retry-After”“302到恶意host”“Content-Length撒谎”。受控 transport 可以精确安排响应并记录请求。测试应断言最终结果、等待秒数、请求头或发送次数等真实边界,而不是断言 mock 对象存在。
对 FastAPI 与 TargetLab 使用 HTTPX ASGITransport,路由、中间件、序列化是真实的,只省去 TCP 端口;这比旧式 TestClient 更接近项目使用的异步 HTTPX。数据库仍是真实 SQLAlchemy schema。
本篇检查点故意放一个坏详情
cd project.\.venv\Scripts\python.exe-m pytest tests\test_pipeline.py::test_crawler_persists_valid_details_and_reports_broken_page-qMapFetcher 按 URL 返回固定列表/详情响应,一条有效、一条缺 title。最终断言有效职位确实写入数据库、坏页进入 errors、任务报告 partial 所需计数正确、快照仍保存。若 Crawler 因一条坏数据回滚全部成功,或者吞掉错误装作 completed,该测试会失败。
端到端旅程为什么跑三次
第一次验证 created=3;第二次验证 ETag 导致 not_modified=3;修改 TargetLabStore 后第三次验证 updated=1、not_modified=2,数据库仍三行。它捕获的不是单个函数,而是 validators 从 Repository 到 HttpFetcher、304 从 FetchResult 到 Crawler 的跨模块链路。
Compose 旅程再多一层真实 Redis、Celery、PostgreSQL、MinIO 和 TCP。它更慢,不应替代每次几秒完成的单测;在交付和部署变更时运行。外部商业网站不进入默认测试,避免网络和内容变化造成假失败,也避免未授权流量。
如何测试重试而不让测试真的等
HttpFetcher 注入 sleeper,测试用 recorder 记录 1 秒而不真的睡;生产默认 asyncio.sleep。这样仍验证计算出的 Retry-After 值,而测试快速。时间、随机数、DNS、外部传输都是适合注入边界的依赖,但不要为了测试把每个纯函数都包成接口。
覆盖率不是验收标准
100% 行覆盖仍可能没有断言租户隔离、重复运行和错误副作用。更好的问题是:把 304 分支改成解析空 body,哪条测试失败?删掉 DNS 检查,哪条失败?删除唯一约束,哪条失败?每个现实变异都应被至少一个行为测试抓住。
本篇完整流水线模块
这次阅读聚焦异常隔离:列表失败、详情下载失败、304、解析/写库失败分别怎样改变 report、快照和事务。它正是各测试层最终汇合的应用服务。
"""把下载、快照、解析和幂等写入组织成一次可报告的采集运行。"""fromdataclassesimportdataclass,fieldfromtypingimportProtocolfrom.concurrencyimportbounded_mapfrom.http_clientimportFetchResultfrom.parsingimportparse_detail,parse_listingfrom.repositoryimportJobRepositoryfrom.snapshotsimportFileSnapshotStoreclassFetcher(Protocol):asyncdeffetch(self,url:str,*,etag:str|None=None,last_modified:str|None=None,)->FetchResult:...@dataclass(frozen=True)classCrawlFailure:url:strmessage:str@dataclassclassCrawlReport:list_pages:int=0discovered:int=0created:int=0updated:int=0unchanged:int=0not_modified:int=0failed:int=0errors:list[CrawlFailure]=field(default_factory=list)classCrawler:"""一次任务的应用服务;单个详情失败不会抹掉其他成功结果。"""def__init__(self,fetcher:Fetcher,repository:JobRepository,snapshots:FileSnapshotStore,*,detail_concurrency:int=4,)->None:self._fetcher=fetcher self._repository=repository self._snapshots=snapshotsifdetail_concurrency<1:raiseValueError("detail_concurrency must be positive")self._detail_concurrency=detail_concurrencyasyncdefrun(self,seed_url:str,*,tenant_id:str,max_pages:int=10)->CrawlReport:ifmax_pages<1:raiseValueError("max_pages must be at least 1")report=CrawlReport()pending=[seed_url]seen_list_pages:set[str]=set()detail_urls:list[str]=[]seen_details:set[str]=set()whilependingandreport.list_pages<max_pages:page_url=pending.pop(0)ifpage_urlinseen_list_pages:continueseen_list_pages.add(page_url)try:response=awaitself._fetcher.fetch(page_url)self._snapshots.save(response.url,response.body,response.headers)listing=parse_listing(response.text,response.url)report.list_pages+=1fordetail_urlinlisting.detail_urls:ifdetail_urlnotinseen_details:seen_details.add(detail_url)detail_urls.append(detail_url)iflisting.next_urlandlisting.next_urlnotinseen_list_pages:pending.append(listing.next_url)exceptExceptionasexc:report.failed+=1report.errors.append(CrawlFailure(page_url,str(exc)))report.discovered=len(detail_urls)requests:list[tuple[str,str|None,str|None]]=[]fordetail_urlindetail_urls:etag,last_modified=self._repository.get_validators(tenant_id,detail_url)requests.append((detail_url,etag,last_modified))asyncdefdownload(request:tuple[str,str|None,str|None])->tuple[str,FetchResult|None,Exception|None]:detail_url,etag,last_modified=requesttry:response=awaitself._fetcher.fetch(detail_url,etag=etag,last_modified=last_modified)returndetail_url,response,NoneexceptExceptionasexc:returndetail_url,None,exc downloaded=awaitbounded_map(requests,download,limit=self._detail_concurrency)fordetail_url,response,download_errorindownloaded:ifdownload_errorisnotNone:report.failed+=1report.errors.append(CrawlFailure(detail_url,str(download_error)))continueassertresponseisnotNonetry:ifresponse.not_modified:report.not_modified+=1continuesnapshot=self._snapshots.save(response.url,response.body,response.headers)item=parse_detail(response.text,response.url)result=self._repository.upsert(tenant_id,item,etag=response.etag,last_modified=response.last_modified,snapshot_id=snapshot.snapshot_id,)self._repository.commit()setattr(report,result.action,getattr(report,result.action)+1)exceptExceptionasexc:self._repository.rollback()report.failed+=1report.errors.append(CrawlFailure(detail_url,str(exc)))returnreport本篇课后练习
- 为解析器、HTTP、Repository、API、进程内旅程、Compose 旅程各写一个“最适合它发现的错误”,不能重复。
- 修改坏详情测试为三条:一条正常、一条下载503、一条缺title,先手算 CrawlReport 再写断言。
- 运行全量 pytest,记录通过数和耗时;再说明为什么这个数字不能替代 Compose 实际旅程。下一篇会启动六个服务并进行持久化恢复演练。