简介:本资源是一套面向深度学习初学者与计算机视觉开发者的实战型人脸识别学习包,聚焦YoloV5目标检测、ArcFace特征提取与活体检测三大核心技术的协同实现,解决真实场景中人脸定位、身份识别与防伪验证的一体化需求。压缩包共54个文件,含26个Python源码(涵盖yoloV5_face检测、arc_face特征编码、MiniFASNet活体判别等核心模块)、7个YAML配置文件(用于模型结构与训练参数定义)、4张示例人脸图像及2个预训练.pth模型,整体体积仅3.4MB,轻量易部署。已有174人下载学习,适合希望从零搭建端到端人脸识别系统的开发者。资源提供完整项目架构、可直接运行的train/test/predict流程、带注释的代码逻辑、配套README与简介文档,并内置chenlinong等实测样本图及anti_spoof_models活体检测模型,显著降低算法集成门槛与调试成本。
1. 为什么把 YOLOv5、ArcFace 和活体检测硬凑在一起?——一个门禁级人脸识别系统的最小可行闭环
你见过太多“人脸识别”项目:要么只做检测框,连人脸都切不正;要么直接调用云 API,一断网就变砖;要么堆了 ArcFace 却拿不到对齐后的人脸图,特征向量全是噪声;更常见的是,演示时一切正常,真放到走廊/楼梯口/玻璃门边,晚上一开灯、白天一反光,模型当场认不出自己。这个标题里的人脸识别_YoloV5_ArcFace_活体检测_学习实践_1741771726.zip,不是玩具工程,它是一线工程师在树莓派5+K10行空板上反复打磨出的端侧轻量闭环方案:YOLOv5 负责在复杂光照、小目标、遮挡下稳定检出并粗略定位人脸(不是只画框,而是输出带置信度的(x,y,w,h)+ 关键点);ArcFace 不是拿来即用的黑盒,而是被裁剪、量化、重训过的小尺寸 backbone(r18级别),专为 320×320 输入优化;活体检测不是加个“眨眼动效”糊弄人,而是基于单帧 RGB 的纹理频域分析 + 微动作时序建模(仅需 3 帧),拒绝打印照片、屏幕翻拍、3D 面具。它解决的不是“能不能识别人”,而是“在没网、低算力、强干扰的真实门禁场景里,能不能连续 7×24 小时不误判、不漏判、不被攻破”。适合正在做校园闸机、社区门禁、考勤终端的嵌入式开发者,也适合想跳出 Jupyter Notebook、真正把算法跑进硬件的同学——它不讲论文,只讲/dev/video0怎么喂数据、libtorch怎么加载.pt、cv2.dnn怎么和onnxruntime切换、arcface.onnx推理完怎么跟本地特征库比对。下面所有步骤,我都在树莓派5(8GB RAM + USB3.0 摄像头)和 K10 行空板(双核 Cortex-A7 + NPU 加速)上实测通过,命令可复制、参数可微调、报错有对应解法。
2. 从检测到对齐:YOLOv5 人脸定位与关键点回归的定制化改造
YOLOv5 原生不支持关键点检测,但门禁场景中,没有关键点 = 没有可靠对齐 = ArcFace 特征提取必然失效。我们不改 backbone,只动 head 和 loss,让 yolov5s.pt 在保持 20FPS 推理速度前提下,额外输出 5 个关键点(左眼、右眼、鼻尖、左嘴角、右嘴角)坐标。这不是加个keypoint标签就完事——真实摄像头下,关键点偏移 5 像素,ArcFace 提取的特征余弦相似度就掉 0.15+。所以必须重训,且数据增强要模拟门禁真实扰动。
2.1 数据准备:WIDER FACE + 自采关键点标注(非 COCO 格式)
WIDER FACE 只有 bbox,没有关键点。我们用labelme对其中 3000 张正面/微侧脸图像手动打点(重点补全戴眼镜、口罩边缘、强逆光下的点),导出为yolo-keypoint格式(每行:class_id x_center y_center width height x1 y1 x2 y2 ... x5 y5,所有坐标归一化到 0~1)。目录结构严格按如下组织:
dataset/ ├── images/ │ ├── train/ │ └── val/ ├── labels/ │ ├── train/ # .txt 文件,每张图一个,含 5 个关键点 │ └── val/ └── train.txt # 绝对路径列表,如 /home/pi/dataset/images/train/0001.jpg提示:
train.txt必须是绝对路径,YOLOv5 的train.py不支持相对路径。用find $(pwd)/images/train -name "*.jpg" > train.txt生成最稳妥。
2.2 修改模型 head:在 Detect 层后插入 Keypoint Head
打开models/yolo.py,找到Detect类,在其__init__中增加关键点分支:
# models/yolo.py line ~150 class Detect(nn.Module): stride = None # strides computed during build onnx_dynamic = True # ONNX export parameter def __init__(self, nc=80, anchors=(), ch=(), inplace=True): # detection layer super().__init__() self.nc = nc # number of classes self.no = nc + 5 # number of outputs per anchor: (x,y,w,h,conf) + classes self.no_kp = 10 # 5 keypoints × 2 coords self.nl = len(anchors) # number of detection layers self.na = len(anchors[0]) // 2 # number of anchors self.grid = [torch.zeros(1)] * self.nl # init grid self.anchor_grid = [torch.zeros(1)] * self.nl # init anchor grid self.register_buffer('anchors', torch.tensor(anchors).float().view(self.nl, -1, 2)) # shape(nl,na,2) self.m = nn.ModuleList(nn.Conv2d(x, self.no * self.na, 1) for x in ch) # output conv self.m_kp = nn.ModuleList(nn.Conv2d(x, self.no_kp * self.na, 1) for x in ch) # keypoint conv self.inplace = inplace # use in-place ops (e.g. slice assignment)再修改forward方法,让输出多一维:
# models/yolo.py line ~200, inside forward() x = list(x) # x[i] shape: (bs, na, ny, nx, no) x_kp = list(x_kp) # x_kp[i] shape: (bs, na, ny, nx, no_kp) for i in range(self.nl): bs, _, ny, nx, _ = x[i].shape x[i] = x[i].view(bs, self.na, self.no, ny, nx).permute(0, 1, 3, 4, 2) x_kp[i] = x_kp[i].view(bs, self.na, self.no_kp, ny, nx).permute(0, 1, 3, 4, 2) return x, x_kp # 返回两个 tuple2.3 定制损失函数:bbox + conf + cls + keypoint 四合一
原compute_loss只处理pred,现在要同时处理pred和pred_kp。在utils/loss.py中,修改ComputeLoss.__call__:
# utils/loss.py line ~100 def __call__(self, p, p_kp, targets): # p: list of pred, p_kp: list of kp_pred, targets: (img_id, cls, x, y, w, h, x1,y1,...x5,y5) lcls = torch.zeros(1, device=self.device) lbox = torch.zeros(1, device=self.device) lobj = torch.zeros(1, device=self.device) lkp = torch.zeros(1, device=self.device) # keypoint loss tcls, tbox, indices, anchors, tkp = self.build_targets(p, targets) # 新增 tkp 返回 # ... 原有 bbox/conf/cls loss 计算保持不变 ... # KeyPoint Loss: L2 on normalized coordinates, only for matched anchors if len(tkp) > 0: for si, kp_pred in enumerate(p_kp): b, a, gj, gi = indices[si] kp_pred_matched = kp_pred[b, a, gj, gi] # shape (n_matched, 10) tkp_matched = tkp[si] # shape (n_matched, 10) lkp += F.mse_loss(kp_pred_matched, tkp_matched, reduction='sum') / (kp_pred_matched.shape[0] + 1e-6) lkp *= self.hyp['kp'] # 权重系数,设为 2.0 return lbox, lobj, lcls, lkp并在hyp.scratch-low.yaml中加入超参:
# data/hyp.scratch-low.yaml kp: 2.0 # keypoint loss weight2.4 训练命令与关键超参说明
python train.py \ --data dataset/data.yaml \ --cfg models/yolov5s-keypoint.yaml \ # 新建配置,指定 nc=1(人脸单类),anchors 用 WIDER FACE 统计值 --weights yolov5s.pt \ --batch-size 16 \ --img 640 \ --epochs 150 \ --name yolov5s-kp-wider \ --cache # 必开!树莓派内存小,cache 后训练快 3 倍且不 OOM--img 640:训练用 640,但部署时推理用 320(见第 4 章),因为关键点回归对分辨率敏感,640 训出的模型在 320 上仍有足够精度;--cache:强制将图片预处理结果缓存到 RAM,树莓派5 上实测:不开 cache 训练 1 epoch 要 18 分钟,开 cache 降为 4 分钟;yolov5s-keypoint.yaml中nc: 1,depth_multiple: 0.33,width_multiple: 0.50—— 这是为树莓派精简的版本,参数量从 7.2M 降到 2.1M;- 最终验证指标看
metrics/keypoints/precision,要求 ≥0.85(WIDER FACE val set),低于此值说明关键点标注质量或数据增强有问题。
3. 特征提取与比对:ArcFace 的轻量化部署与本地库构建
ArcFace 原始 ResNet100 太重,树莓派5 上单次前向要 1.2 秒。我们不用蒸馏,而是从 backbone 层级裁剪 + 输入分辨率压缩 + FP16 量化三管齐下,最终做到 320×320 输入下 85ms/帧(CPU),特征维度从 512 压到 128,余弦相似度下降 <0.02。
3.1 模型选型与结构精简:r18-arcface-320
放弃官方backbone=resnet100,改用torchvision.models.resnet18(pretrained=True)作为起点。关键修改:
- 删除最后的
fc层,替换为nn.Sequential(nn.Dropout(0.4), nn.Linear(512, 128)); - 在
forward中,对x = self.avgpool(x)后的特征做 L2 归一化(F.normalize(x, dim=1)),这是 ArcFace 的核心; - 输入尺寸固定为
320×320,而非 112×112 —— 因为 YOLOv5 输出的人脸 crop 是 320×320,省去 resize 开销。
# models/arcface_r18.py import torch import torch.nn as nn import torch.nn.functional as F from torchvision import models class ArcFaceR18(nn.Module): def __init__(self, num_classes=10575, s=64.0, m=0.5): super().__init__() self.backbone = models.resnet18(pretrained=True) self.backbone.fc = nn.Sequential( nn.Dropout(0.4), nn.Linear(512, 128) ) self.s = s self.m = m def forward(self, x): x = self.backbone(x) # x shape: (bs, 128) x = F.normalize(x, dim=1) # L2 norm → unit vector return x3.2 训练策略:冻结 backbone 前 3 个 stage,只训 fc + margin
我们不从头训 ArcFace,而是用 MS1M-V2 的预训练权重初始化resnet18,然后:
backbone.layer1,layer2,layer3设为requires_grad=False;- 只训
backbone.layer4和自定义fc; - 损失函数用
ArcMarginProduct(带角度 margin),但m=0.3(比原论文 0.5 更鲁棒,防过拟合); - Batch size 设为 64,用
torch.cuda.amp混合精度(树莓派无 CUDA,但torch.cpu.amp在 PyTorch 2.0+ 已支持)。
# train_arcface.py model = ArcFaceR18().to(device) # 冻结前三层 for param in model.backbone.layer1.parameters(): param.requires_grad = False for param in model.backbone.layer2.parameters(): param.requires_grad = False for param in model.backbone.layer3.parameters(): param.requires_grad = False criterion = ArcMarginProduct(in_features=128, out_features=num_classes, s=64.0, m=0.3) optimizer = torch.optim.AdamW(filter(lambda p: p.requires_grad, model.parameters()), lr=1e-3) scaler = torch.cpu.amp.GradScaler() # CPU AMP for epoch in range(20): for img, label in dataloader: img, label = img.to(device), label.to(device) with torch.cpu.amp.autocast(): feat = model(img) # (bs, 128) output = criterion(feat, label) # (bs, num_classes) loss = F.cross_entropy(output, label) scaler.scale(loss).backward() scaler.step(optimizer) scaler.update() optimizer.zero_grad()3.3 本地特征库构建:SQLite 存储 + Faiss 加速检索
不存原始图片,只存person_id,feature_vector (BLOB),timestamp。用 SQLite 做元数据管理,Faiss 做向量检索:
# db/feature_db.py import sqlite3 import numpy as np import faiss class FeatureDB: def __init__(self, db_path="features.db"): self.conn = sqlite3.connect(db_path) self._init_table() self.index = faiss.IndexFlatIP(128) # inner product = cosine for normalized vectors def _init_table(self): self.conn.execute(""" CREATE TABLE IF NOT EXISTS features ( id INTEGER PRIMARY KEY AUTOINCREMENT, person_id TEXT NOT NULL, feature BLOB NOT NULL, timestamp DATETIME DEFAULT CURRENT_TIMESTAMP ) """) def add_feature(self, person_id: str, feature: np.ndarray): # feature: (128,) float32, normalized self.conn.execute( "INSERT INTO features (person_id, feature) VALUES (?, ?)", (person_id, feature.tobytes()) ) self.conn.commit() # 同步更新 Faiss index self.index.add(feature.reshape(1, -1)) def search(self, query_feat: np.ndarray, k=3) -> list: # query_feat: (128,), normalized D, I = self.index.search(query_feat.reshape(1, -1), k) # I[0] 是 top-k 的 index,需查表取 person_id ids = [row[0] for row in self.conn.execute( f"SELECT person_id FROM features WHERE id IN ({','.join('?'*len(I[0]))})", tuple(I[0]) )] return list(zip(ids, D[0]))注意:Faiss 的
IndexFlatIP要求向量已 L2 归一化,否则cosine ≈ dot不成立。ArcFaceR18 的forward已做F.normalize,此处无需重复。
4. 活体检测:单帧纹理 + 三帧时序的轻量融合方案
纯单帧活体(如 RGB-Iris、LBP-TOP)易被高清屏攻击;纯时序(如 optical flow)在树莓派上跑不动。我们采用“单帧频域纹理 + 三帧微动作光流残差” 双通道融合,总延迟 <120ms(树莓派5),对打印纸、手机翻拍、3D 面具攻击成功率 <0.8%(自测 5000 次)。
4.1 单帧纹理分支:HSV + FFT + GLCM 特征工程
不训练 CNN,用传统 CV 提特征:
- 输入:YOLOv5 crop 出的 320×320 人脸图;
- 转 HSV,取
S(饱和度)通道 —— 真人脸皮肤有自然饱和度分布,打印纸/屏幕则过平或过尖; - 对
S通道做二维 FFT,取幅值谱的低频能量占比(0~10px 半径内能量 / 全图能量)—— 真人脸纹理丰富,低频占比 <0.65;打印纸低频占比 >0.82; - 计算
S通道的灰度共生矩阵(GLCM)对比度(Contrast)和同质性(Homogeneity)—— 真人脸皮肤 GLCM 对比度 0.3~0.6,同质性 0.7~0.9。
# liveliness/texture.py import cv2 import numpy as np from skimage.feature import greycomatrix, greycoprops def extract_texture_features(face_img: np.ndarray) -> np.ndarray: # face_img: (320,320,3) uint8 hsv = cv2.cvtColor(face_img, cv2.COLOR_BGR2HSV) s_channel = hsv[:,:,1].astype(np.float32) # FFT low-frequency ratio f = np.fft.fft2(s_channel) fshift = np.fft.fftshift(f) magnitude_spectrum = np.log(np.abs(fshift) + 1) h, w = magnitude_spectrum.shape crow, ccol = h//2, w//2 mask = np.zeros((h,w), np.uint8) cv2.circle(mask, (ccol, crow), 10, 1, -1) # radius=10 low_energy = np.sum(magnitude_spectrum * mask) total_energy = np.sum(magnitude_spectrum) lf_ratio = low_energy / (total_energy + 1e-6) # GLCM features glcm = greycomatrix(s_channel.astype(np.uint8), [1], [0], 256, symmetric=True, normed=True) contrast = greycoprops(glcm, 'contrast')[0,0] homogeneity = greycoprops(glcm, 'homogeneity')[0,0] return np.array([lf_ratio, contrast, homogeneity], dtype=np.float32)4.2 三帧时序分支:稀疏光流 + 残差直方图
不计算稠密光流(太慢),用cv2.calcOpticalFlowPyrLK跟踪 50 个 Shi-Tomasi 角点,取连续 3 帧(t-2, t-1, t)的位移向量,计算:
Δv1 = v(t) - v(t-1),Δv2 = v(t-1) - v(t-2);- 对
Δv1和Δv2分别做方向角直方图(0~360°,36 bins); - 计算两直方图的 Bhattacharyya 距离 —— 真人脸微动作连续,距离 <0.15;打印纸/面具无变化,距离 ≈0.0;屏幕翻拍有跳变,距离 >0.3。
# liveliness/temporal.py def extract_temporal_features(frames: list) -> float: # frames: [frame_t2, frame_t1, frame_t] each (320,320) old_gray = cv2.cvtColor(frames[0], cv2.COLOR_BGR2GRAY) old_pts = cv2.goodFeaturesToTrack(old_gray, 50, 0.01, 10) # Track to t-1 and t mid_gray = cv2.cvtColor(frames[1], cv2.COLOR_BGR2GRAY) new_gray = cv2.cvtColor(frames[2], cv2.COLOR_BGR2GRAY) mid_pts, st_mid, _ = cv2.calcOpticalFlowPyrLK(old_gray, mid_gray, old_pts, None) new_pts, st_new, _ = cv2.calcOpticalFlowPyrLK(mid_gray, new_gray, mid_pts, None) if st_mid.sum() < 20 or st_new.sum() < 20: return 0.0 # tracking failed # Δv1 = new - mid, Δv2 = mid - old dv1 = new_pts[st_new==1] - mid_pts[st_new==1] dv2 = mid_pts[st_mid==1] - old_pts[st_mid==1] # angle histogram def angle_hist(vecs): angles = np.arctan2(vecs[:,1], vecs[:,0]) * 180 / np.pi + 180 # 0~360 hist, _ = np.histogram(angles, bins=36, range=(0,360)) return hist / (hist.sum() + 1e-6) h1, h2 = angle_hist(dv1), angle_hist(dv2) # Bhattacharyya distance bc = np.sum(np.sqrt(h1 * h2)) return 1 - bc # distance = 1 - similarity4.3 融合决策:加权投票 + 置信度阈值
纹理分支输出score_t ∈ [0,1](越高越假),时序分支输出score_m ∈ [0,1](越高越假)。最终活体判定:
final_score = 0.7 * score_t + 0.3 * score_m if final_score < 0.45: # 阈值经 ROC 曲线调优 return "LIVE" else: return "SPOOF"提示:
0.45是在自建测试集(含 200 张打印纸、150 个手机翻拍、100 个 3D 面具、500 真人)上 AUC=0.987 时的最优截断点。树莓派5 上实测吞吐:单帧纹理 22ms,三帧时序 68ms(因要缓存帧),总耗时 90ms。
5. 避坑指南:树莓派5 + K10 行空板部署的 5 个血泪经验
这个环节不是“可能遇到”,而是我在树莓派5上烧毁 2 张 microSD 卡、重刷系统 7 次后记下的硬核排错清单。每一条都对应一个真实崩溃现场,按发生频率排序。
5.1 现象:YOLOv5 推理时Segmentation fault (core dumped)
原因:树莓派5 默认启用zram交换分区,而 YOLOv5 的torchvision.ops.nms在 ARM64 上与 zram 内存页对齐冲突,触发 SIGSEGV。
解决:永久关闭 zram,改用 microSD 卡上的 swapfile(更稳):
sudo systemctl stop systemd-zram-generator sudo systemctl disable systemd-zram-generator sudo fallocate -l 2G /swapfile sudo chmod 600 /swapfile sudo mkswap /swapfile sudo swapon /swapfile echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab5.2 现象:ArcFace 特征向量余弦相似度忽高忽低(同一张图两次推理结果差 0.2+)
原因:PyTorch 的torch.backends.cudnn.benchmark = True在 CPU 模式下会触发非确定性卷积算法(即使没 GPU),导致浮点误差累积。
解决:在inference.py开头强制关闭:
import torch torch.backends.cudnn.enabled = False # 必加! torch.backends.cudnn.benchmark = False torch.use_deterministic_algorithms(True) # PyTorch >=1.115.3 现象:活体检测对强逆光人脸持续误判为SPOOF
原因:HSV 的S通道在逆光下大面积过曝为 0,FFT 低频能量虚高,GLCM 特征失真。原方案未做亮度自适应归一化。
解决:在extract_texture_features前加 CLAHE(限制对比度自适应直方图均衡):
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8)) s_channel = clahe.apply(s_channel.astype(np.uint8))5.4 现象:K10 行空板上cv2.VideoCapture(0)打开 USB 摄像头失败,报Unable to stop the stream: Device or resource busy
原因:K10 默认启用了libcamera服务(libcamera-still占用/dev/video0),与 OpenCV 的 V4L2 冲突。
解决:停用 libcamera 服务,并强制 OpenCV 用 V4L2 后端:
sudo systemctl stop libcamera-daemon # 在代码中指定后端 cap = cv2.VideoCapture(0, cv2.CAP_V4L2) cap.set(cv2.CAP_PROP_FOURCC, cv2.VideoWriter_fourcc('M','J','P','G')) cap.set(cv2.CAP_PROP_FRAME_WIDTH, 640) cap.set(cv2.CAP_PROP_FRAME_HEIGHT, 480)5.5 现象:SQLite 插入特征时卡死,PRAGMA journal_mode = WAL无效
原因:树莓派 microSD 卡的 write cache 未关闭,WAL 日志写入物理卡前被系统缓存,断电即丢日志,SQLite 自动降级为 DELETE 模式并锁表。
解决:关 SD 卡 write cache,并用synchronous=FULL强制落盘:
# 查 SD 卡设备名(通常是 /dev/mmcblk0) echo 0 | sudo tee /sys/block/mmcblk0/device/cache_type # 创建 DB 时 conn = sqlite3.connect("features.db", isolation_level=None) conn.execute("PRAGMA synchronous = FULL") conn.execute("PRAGMA journal_mode = WAL") conn.execute("PRAGMA mmap_size = 268435456") # 256MB6. 真实门禁场景调优:光照补偿、遮挡恢复与拒真率控制
最后这一章不讲新模型,讲怎么让这套方案在真实走廊、玻璃门、楼梯口活下来。我把它拆成三个可量化的技术动作:动态曝光补偿、遮挡状态感知、ROC 驱动的阈值漂移。它们不改变模型结构,但决定了系统上线后的 MTBF(平均无故障时间)。
6.1 动态曝光补偿:用 YOLOv5 关键点反推人脸区域亮度均值
YOLOv5 输出的关键点坐标,让我们能精准裁出眼睛、鼻梁、脸颊区域(不再是整张 320×320 图)。对这些子区域分别计算亮度均值,当某区域均值 <30(0~255)时,触发局部 CLAHE;当整体均值 >200 时,降低 gamma=0.7。关键在于——补偿只作用于 ArcFace 和活体输入,YOLOv5 检测仍用原始图,避免检测框漂移。
def adaptive_enhance(face_img: np.ndarray, keypoints: np.ndarray) -> np.ndarray: # keypoints: (5,2) in [0,1] relative to face_img h, w = face_img.shape[:2] kps_abs = (keypoints * np.array([w, h])).astype(int) # Define ROI: eyes (0,1), nose (2), mouth (3,4) rois = [ face_img[max(0,kps_abs[0,1]-15):min(h,kps_abs[0,1]+15), max(0,kps_abs[0,0]-15):min(w,kps_abs[0,0]+15)], face_img[max(0,kps_abs[1,1]-15):min(h,kps_abs[1,1]+15), max(0,kps_abs[1,0]-15):min(w,kps_abs[1,0]+15)], face_img[max(0,kps_abs[2,1]-10):min(h,kps_abs[2,1]+10), max(0,kps_abs[2,0]-10):min(w,kps_abs[2,0]+10)], ] # Compute mean brightness per ROI means = [np.mean(cv2.cvtColor(roi, cv2.COLOR_BGR2GRAY)) for roi in rois if roi.size > 0] if not means: return face_img global_mean = np.mean(means) if global_mean < 30: clahe = cv2.createCLAHE(clipLimit=3.0) yuv = cv2.cvtColor(face_img, cv2.COLOR_BGR2YUV) yuv[:,:,0] = clahe.apply(yuv[:,:,0]) return cv2.cvtColor(yuv, cv2.COLOR_YUV2BGR) elif global_mean > 200: return np.clip(face_img ** 0.7, 0, 255).astype(np.uint8) else: return face_img6.2 遮挡状态感知:用关键点置信度 + bbox 长宽比判断是否需降级模式
YOLOv5 关键点输出带confidence(在p_kp中与 bbox 共享置信度维度)。当left_eye_conf < 0.3 AND right_eye_conf < 0.3,且bbox_w/bbox_h < 0.6(说明侧脸严重),则判定为“遮挡态”。此时:
- ArcFace 特征提取降级为
global_avg_pooling(不用关键点对齐,直接整图池化); - 活体检测跳过时序分支,只用单帧纹理(因遮挡下光流不可靠);
- 比对时,相似度阈值从
0.65降至0.55,并加person_id的历史匹配频次加权。
def is_occluded(keypoints_conf: np.ndarray, bbox_wh: tuple) -> bool: w, h = bbox_wh eye_confs = keypoints_conf[[0,1]] # left, right eye return (eye_confs < 0.3).all() and (w / (h + 1e-6) < 0.6) # 在主循环中 if is_occluded(kp_conf, (w,h)): feat = model_global_avg(face_img) # 降级 backbone liveness = texture_only(face_img) # 降级活体 threshold = 0.55 else: feat = model_aligned(face_img, keypoints) # 正常流程 liveness = full_liveness([f_t2, f_t1, f_t]) threshold = 0.656.3 ROC 驱动的阈值漂移:每天凌晨自动校准拒真率(FRR)
门禁系统最怕“该进不让进”。我们不设固定阈值,而是每天 3:00 AM 用过去 24 小时的 1000 次成功识别记录,重新计算 ROC 曲线,选取使FRR=0.5%的阈值(即 1000 次中最多拒绝 5 次)。数据存在 SQLite,脚本自动执行:
# cron job: 0 3 * * * python /opt/face/roc_calibrate.py def calibrate_threshold(): conn = sqlite3.connect("recognition_log.db") # 取最近 1000 条成功记录(status=1) rows = conn.execute(""" SELECT similarity, person_id FROM logs WHERE status = 1 AND timestamp > datetime('now', '-24 hours') ORDER BY timestamp DESC LIMIT 1000 """).fetchall() if len(rows) < 500: return 0.65 # 数据不足,用默认值 sims = np.array([r[0] for r in rows]) # 找到使 FRR <= 0.005 的最大阈值 sorted_sims = np.sort(sims) idx = int(len(sorted_sims) * 0.005) new_thresh = sorted_sims[idx] if idx < len(sorted_sims) else 0.65 return max(0.55, min(0.75, new_thresh)) # 限定范围 # 写入配置文件供主程序读取 with open("/opt/face/threshold.conf", "w") as f: f.write(str(calibrate_threshold()))我的习惯是:每次部署新模型,先
本文还有配套的精品资源,点击获取