1. 为什么我要把 100 道 LLM 面试选择题搬进本地脚本
LLM 面试选择题题库,指的是把 Transformer 基础、模型架构、训练微调、推理优化、工程部署这五大领域的单选题整理成结构化数据,再用脚本批量调用大模型核对答案与解析。它适合正在准备算法岗面试、需要反复刷题又不想手动对答案的开发者,也适合想用一套统一 Key 同时跑通多个模型做交叉验证的人。
我手上这份题库有 100 道题,覆盖 Q01 到 Q100,每题包含题干、四个选项、标准答案和一段解析。手动刷一遍大概 1.5 小时,但真正麻烦的是核对环节:同一道题我想知道不同模型给出的答案是否一致,如果每个模型都单独申请 Key、单独改环境变量,光是配置就能耗掉半小时。更别说有些题涉及 RoPE、GQA、PagedAttention 这些细节,单模型答案未必可靠,多模型交叉验证才更稳。
所以这篇的做法是:把题库存成本地 JSON,写一个刷题脚本,用 TaoToken 作为统一的 Key 和 API 通道,一次性接入多个模型。你只需要维护一份 config.toml 和一份 settings.json,脚本里切换模型只改一个字段。下面从环境准备讲到可复制的请求验证,最后给出我踩过的几个坑。
2. TaoToken 前置:统一 Key 与 API 通道怎么理解
TaoToken 在这里扮演的角色,是一个统一的模型调用入口。你可以把它理解成“一个 Key 走通多个模型”的通道:脚本不需要为每个模型维护不同的 base_url 和鉴权方式,只要把请求发到同一个 API 地址,在 body 里指定模型名即可。对于刷题这种需要频繁切换模型核对答案的场景,省掉的是重复配置成本。
需要提前准备的东西只有两样:一个可用的 API Key,以及确认你要调用的模型名。Key 在控制台的 API Keys 页面创建,模型名以文档里列出的为准。官网地址是 https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= ,API 基址是 https://taotoken.net/api ,注意 API 地址不带 UTM 参数,配置里直接写这个。
注意:API Key 不要写进会被 git 跟踪的文件。我习惯把它放在环境变量里,config.toml 只引用变量名,settings.json 里也不出现明文。
创建 Key 的入口在控制台:https://taotoken.net/console?utm_source=taotoken_aicg_blog_end&utm_content=console&utm_campaign=rewrite ,Key 管理页在 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite 。如果你只是想先手动验证某道题的答案,可以直接用模型对话页试:https://taotoken.net/models?utm_source=taotoken_aicg_blog_end&utm_content=model_chat&utm_campaign=rewrite 。
3. 可复制配置:config.toml 与 settings.json 骨架
先建目录结构,我习惯这样放:
llm-quiz/ ├── config.toml ├── settings.json ├── questions.json └── quiz_runner.pyconfig.toml 负责通道级配置,settings.json 负责刷题行为配置。两者分离的好处是:换模型只动 settings.json,换通道只动 config.toml。
config.toml 骨架如下,重点是 base_url 指向 TaoToken 的 API 地址,api_key 从环境变量读取:
# config.toml [provider] name = "taotoken" base_url = "https://taotoken.net/api" api_key_env = "TAOTOKEN_API_KEY" timeout_seconds = 60 max_retries = 3 [models] # 用于交叉验证的模型列表,按需增删 candidates = [ "claude-3-5-sonnet", "gpt-4o", "deepseek-chat" ] [request] temperature = 0.0 max_tokens = 512settings.json 骨架如下,控制刷题范围、限时和输出格式:
{ "quiz": { "question_file": "questions.json", "domains": ["transformer", "architecture", "training", "inference", "engineering"], "limit_per_question_seconds": 90, "shuffle": false }, "verify": { "enabled": true, "models": ["claude-3-5-sonnet", "gpt-4o"], "require_consensus": true, "save_disagreements": "disagreements.json" }, "output": { "report_file": "report.md", "show_explanation": true } }questions.json 的结构建议统一成下面这样,方便脚本解析。题干、选项、答案、解析四个字段是必须的:
[ { "id": "Q01", "domain": "transformer", "stem": "在Scaled Dot-Product Attention中,除以√dₖ的主要目的是什么?", "options": { "A": "减少计算量", "B": "防止过拟合", "C": "避免softmax梯度消失", "D": "增加模型容量" }, "answer": "C", "explanation": "dₖ较大时,QKᵀ的方差为dₖ,softmax会进入梯度饱和区。除以√dₖ将方差缩放到1,保持梯度敏感。" } ]把 100 道题按这个结构录入后,脚本就能逐题读取。录入时注意选项键统一用大写字母,答案字段只存字母,解析单独存,这样后面做多模型比对时不会因为格式差异误判。
4. 刷题脚本:把 TaoToken 接进请求链路
脚本用 Python 写,依赖只有 requests。核心逻辑是:读 config.toml 拿通道信息,读 settings.json 拿模型列表,逐题构造 prompt 发给 TaoToken,收集各模型答案后与标准答案比对。
# quiz_runner.py import json import os import time import tomllib import requests def load_config(path="config.toml"): with open(path, "rb") as f: return tomllib.load(f) def load_settings(path="settings.json"): with open(path, "r", encoding="utf-8") as f: return json.load(f) def build_prompt(q): opts = "\n".join(f"{k}. {v}" for k, v in q["options"].items()) return ( f"题目:{q['stem']}\n{opts}\n" "请只输出一个选项字母(A/B/C/D),不要输出其他内容。" ) def ask_model(cfg, model, prompt): api_key = os.environ[cfg["provider"]["api_key_env"]] url = cfg["provider"]["base_url"].rstrip("/") + "/v1/chat/completions" headers = { "Authorization": f"Bearer {api_key}", "Content-Type": "application/json", } payload = { "model": model, "messages": [{"role": "user", "content": prompt}], "temperature": cfg["request"]["temperature"], "max_tokens": cfg["request"]["max_tokens"], } for attempt in range(cfg["provider"]["max_retries"]): try: resp = requests.post( url, headers=headers, json=payload, timeout=cfg["provider"]["timeout_seconds"], ) resp.raise_for_status() return resp.json()["choices"][0]["message"]["content"].strip() except Exception as e: if attempt == cfg["provider"]["max_retries"] - 1: return f"ERROR: {e}" time.sleep(1.5 * (attempt + 1)) def run(): cfg = load_config() st = load_settings() questions = json.load(open(st["quiz"]["question_file"], encoding="utf-8")) disagreements = [] report = [] for q in questions: prompt = build_prompt(q) answers = {} for model in st["verify"]["models"]: answers[model] = ask_model(cfg, model, prompt) consensus = len(set(answers.values())) == 1 hit = answers.get(st["verify"]["models"][0]) == q["answer"] report.append({ "id": q["id"], "standard": q["answer"], "answers": answers, "consensus": consensus, "first_model_hit": hit, }) if not consensus: disagreements.append(q["id"]) json.dump(report, open(st["output"]["report_file"], "w", encoding="utf-8"), ensure_ascii=False, indent=2) json.dump(disagreements, open(st["verify"]["save_disagreements"], "w", encoding="utf-8"), ensure_ascii=False, indent=2) print(f"完成 {len(questions)} 题,分歧题 {len(disagreements)} 道") if __name__ == "__main__": run()运行前设置环境变量:
export TAOTOKEN_API_KEY="你的Key" python quiz_runner.py脚本里有两个设计点值得说明。第一,temperature 设为 0.0,选择题需要确定性输出,避免同一题两次调用答案不同。第二,max_tokens 设 512 足够,因为 prompt 明确要求只输出一个字母,模型不会长篇大论。如果你发现某些模型不遵守“只输出字母”的约束,可以在解析时加一层正则提取,但优先用 prompt 约束。
5. 验证请求:一次可复制的成功结果
在跑全量 100 题之前,先用一条最小请求确认链路通。下面这条 curl 可以直接复制,把 Key 换成你自己的:
curl -s https://taotoken.net/api/v1/chat/completions \ -H "Authorization: Bearer $TAOTOKEN_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-3-5-sonnet", "messages": [ {"role": "user", "content": "在Scaled Dot-Product Attention中,除以√dₖ的主要目的是什么?A.减少计算量 B.防止过拟合 C.避免softmax梯度消失 D.增加模型容量。只输出字母。"} ], "temperature": 0.0, "max_tokens": 16 }'预期返回结构里,choices[0].message.content 应该是 "C"。如果返回的是完整句子,说明模型没遵守约束,但链路本身是通的。我实测下来,这条请求在正常网络下 2 到 4 秒返回,返回体里 usage 字段能看到 prompt_tokens 和 completion_tokens,方便你估算 100 题的总消耗。
确认单条通之后,再跑脚本。第一次跑建议把 settings.json 里的 models 只留一个,先确认 100 题能完整跑完不中断,再加第二个模型做交叉验证。全量跑完的 report.md 里,每道题会记录标准答案、各模型答案、是否一致、首个模型是否命中。分歧题单独存到 disagreements.json,这些题往往就是题库里最容易混淆的知识点,值得重点看解析。
6. 本篇常见错排查
报错 401 Unauthorized:九成是环境变量没生效。先echo $TAOTOKEN_API_KEY确认有值,再检查 config.toml 里的 api_key_env 名字和实际导出的变量名是否一致。注意不要在 config.toml 里直接写 Key 明文,容易在复制时带上多余空格。
报错 404 或路径不对:base_url 写成了带 /v1 的完整路径,脚本里又拼了一次 /v1/chat/completions,导致重复。config.toml 里 base_url 只写到 https://taotoken.net/api ,路径拼接交给脚本。
模型名不存在:settings.json 里的模型名必须和文档里列出的完全一致,大小写和连字符都不能错。不确定时先用模型对话页手动选一次,确认能出结果再写进配置。
脚本卡住不返回:timeout_seconds 设太短,或者某道题 prompt 过长。100 道题里工程部署部分有些题干较长,60 秒通常够,如果网络波动可以调到 90。max_retries 设 3 次,配合退避重试能覆盖大部分瞬时失败。
多模型答案全不一致:先看是不是 prompt 里选项顺序被模型重排了。我的做法是选项固定用 A/B/C/D 标注,prompt 里明确“只输出一个选项字母”,如果还有模型输出“答案是C”这种,就在解析层加正则[ABCD]提取首个匹配。
分歧题太多:如果 disagreements.json 里超过 20 道,说明模型对某些领域确实存在认知差异,这时候不要盲目相信标准答案,把分歧题单独拎出来,用模型对话页逐题追问解析,人工判断哪个更合理。这也是多模型交叉验证的价值所在。
7. 继续往下走:从刷题到长期编码练习
跑通这套本地刷题环境后,你手上就有了一个可复用的多模型调用骨架。同样的 config.toml 和请求封装,可以直接拿去做其他需要批量核对的任务,比如把面试题换成代码题、把选择题换成判断题。如果你打算长期做这类练习,甚至把模型接进日常编码流程里做辅助,可以了解一下 Coding Plan:https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding_plan&utm_campaign=rewrite ,它更适合需要持续调用、按周期使用的场景。
接入细节和参数说明以官方文档为准:https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite 。如果你用的是 Claude Code 这类工具,Anthropic 兼容接入的说明在这里:https://taotoken.net/claude-code-anthropic?utm_source=taotoken_aicg_blog_end&utm_content=claudecode&utm_campaign=rewrite 。先把 100 题跑一遍,把分歧题吃透,比刷十遍标准答案更有用。