1.4 加入单元测试结果和高清下载选项

This commit is contained in:
siknet
2026-09-19 22:10:41 +08:00
parent e2c643b6be
commit 0d91ec7c2a
17 changed files with 457 additions and 256 deletions
+32
View File
@@ -0,0 +1,32 @@
name: CI Tests
on:
push:
branches: [ main, master ]
pull_request:
branches: [ main, master ]
jobs:
test:
runs-on: ubuntu-latest
strategy:
matrix:
python-version: ["3.9", "3.10", "3.11", "3.12"]
steps:
- uses: actions/checkout@v4
- name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt
pip install -r requirements-dev.txt
- name: Run pytest
run: |
pytest -v tests/
+2
View File
@@ -0,0 +1,2 @@
/build
/dist
@@ -1,209 +0,0 @@
# FreePEP 📚 人教社中小学电子教材批量下载器
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
**FreePEP** 是一款专为[人民教育出版社中小学电子教材平台](https://jc.pep.com.cn/)开发的自动化教材解析、批量抓取与高清 PDF 合成工具。
提供**WebUI 界面**与**交互式命令行**,内置全量 780+ 本教材目录(数据截止到2026年08月31日)自动解密引擎与阿里云 WAF 滑块验证码自动破解机制,支持一键下载指定学段、学科、年级的全套教材并自动生成高清 PDF 文件。
**更新内容:**
2026.09.04 FreePEP v1.1 命令行加入了批量下载功能,详情见**启动办法三**
---
## ✨ 核心特性
- 🎯 **双操作模式**:
- **现代化 WebUI 界面**:全响应式布局,还原官网层级筛选体验,支持复选框一键批量下载、实时进度条展示与本地保存目录快捷打开。
- **交互式 CLI 终端**:支持数字菜单导航选择与命令行参数直达(脚本集成与服务器环境友好)。
- 📦 **高清单页下载与 PDF 合成**:自动探测各教材实际总页数,批量下载高分辨率原始 JPG,并通过 Pillow 库自动合成为标准 PDF 文件。
---
## 📸 界面预览
### WebUI 网页端
![](https://img.wemd.app/1788139668274_4n5d28.png)
---
### CLI 网页端
![](https://img.wemd.app/1788139734008_2tmuak.png)
# 🚀 使用方式
## 🖥️ 方式一:(推荐,适合小白)
直接下载发行包,解压缩后运行FreePEP.exe,在弹出的Web页面里面自行操作下载。
## 方式二:从源码启动
### 克隆仓库与安装依赖
```bash
# 克隆仓库
git clone https://github.com/siknet/FreePEP.git
cd FreePEP
# 安装 Python 依赖
pip install -r requirements.txt
# 安装 Playwright 所需的 Chromium 浏览器内核
playwright install chromium
```
### 启动办法一:Webui
在终端中执行以下命令,系统将自动启动本地服务并在默认浏览器中打开管理页面:
```bash
python webui.py
```
- **访问地址**:`http://127.0.0.1:8000`
- **默认下载目录**:项目根目录下的 `./downloads` 文件夹(可在 Web 界面点击「📁 打开保存目录」直达)。
---
### 启动办法二:使用 CLI 命令行
#### 1. 交互式菜单模式
直接运行 `cli.py`,根据控制台提示逐步选择学段、学科与年级:
```bash
python cli.py
```
#### 2. 参数直达模式(适合自动化与脚本调用)
通过命令行参数直接指定条件过滤并自动开始下载:
```bash
# 1. 下载小学一年级的所有学科教材
python cli.py --xd "小学" --nj "一年级" -y
# 2. 下载高中数学的所有必修/选修教材
python cli.py --xd "高中" --xk "数学" -y
# 3. 下载初中全部教材
python cli.py --xd "初中" -y
# 4. 全局关键词搜索并下载(如包含"物理"的所有教材)
python cli.py --search "物理"
# 5. 清理本地临时图片缓存
python cli.py --clear-cache
# 6. 自定义 PDF 输出路径
python cli.py --xd "小学(六三学制)" --xk "语文" --nj "一年级" -o "D:/Textbooks" -y
```
**参数说明**:
| 参数 | 说明 | 示例 |
| :--- | :--- | :--- |
| `--xd` | 指定学段 | `小学(六三学制)`、`初中(六三学制)`、`小学(五四学制)`、`初中(五·四学制)`、`高中`、`培智学校`、`聋校`、`盲校(盲文版)`、`盲校(低视力版)` |
| `--xk` | 指定学科 | `语文`、`数学`、`英语`、`物理`、`化学`、`历史`、`道德与法治` 等 |
| `--nj` | 指定年级/册次 | `一年级`、`二年级` ... `九年级`、`专项`、`必修` 等 |
| `--search`, `-s` | 关键词全局搜索 | `必修`、`高一`、`地理` |
| `--clear-cache`, `-c` | 清理本地下载临时图片缓存 (`temp_pages`) | 无需参数 |
| `--refresh`, `-r` | 强制重新从官方服务器拉取解密最新目录 | 无需参数 |
| `--output`, `-o` | 指定 PDF 保存目录 | 默认: `./downloads` |
| `--yes`, `-y` | 跳过确认提示直接开始下载 | 开启免交互 |
### 启动办法三:按「学段 ➔ 年级」层级批量多线程下载 (download_all.py)
如果您希望将教材按照 **`学段/年级/`** 两层规范文件夹分类归档下载,直接运行专属的高速批量多线程下载脚本:
```bash
# 1. 一键下载全网全量教材(默认 3 线程并发,按「学段/年级」两层子目录自动归档)
python download_all.py
# 2. 启用多线程并发加速(例如 10 线程并行高速下载)
python download_all.py -w 10
# 3. 下载多个指定学段(支持逗号分隔,如:义务教育六三学制、五四学制与高中)并自定义保存路径
python download_all.py --xd "义务教育(六三学制),义务教育(五四学制),高中" -w 10 -o "D:/人教社教材"
# 4. 仅下载单个学段(如高中全部教材)
python download_all.py --xd "高中"
# 5. 仅下载某个指定年级教材
python download_all.py --nj "一年级"
```
**功能特性**:
- 🚀 **多线程并发提速**:默认 **3 线程**并发同时下载,支持通过 `-w / --workers` 自定义并发线程数(如 5~10 线程),下载效率大幅提升。
- 🎯 **支持多学段组合**:`--xd` 支持传入多个学段(逗号分隔),精准下载普通义务教育或高中学段,自动跳过特教教材。
- 🌲 **两层层级目录**:自动分类保存为 `downloads/<学段>/<年级>/<教材名>.pdf`,告别成百上千文件堆在单目录。
- ⚡ **智能断点续传**:已完整下载的教材**自动秒跳过**,中途随时中断无缝继续,绝不重复下载。
- 🧹 **切片自动清理**:每本教材合成 PDF 后自动清理临时图片分片,极大保护硬盘空间。
- 💬 **清爽紧凑输出**:不打印繁琐单页过程,仅在教材下载合并完成或跳过时输出进度与提示。
---
## 📦 一键打包为 Windows 独立 EXE 发行版
如果您想将本项目打包成 **脱离 Python 环境** 的独立绿色软件包发布给普通用户,只需执行工作区自带的一键打包脚本:
```bash
python build_exe.py
```
### 打包脚本会自动完成以下操作:
1. 自动检查并安装 `PyInstaller` 编译工具。
2. 将 WebUI 及所有 Python 依赖打包为独立可执行文件 `FreePEP.exe`。
3. **自动提取并内嵌绿色便携版 Chromium 浏览器内核**至 `browsers/` 目录。
4. 自动在 `dist/` 目录下生成 `FreePEP-Windows-x64.zip` 发行压缩包。
> **分发给用户使用**:用户下载压缩包后解压,**双击 `FreePEP.exe` 即可直接使用**(会自动弹出系统默认浏览器打开 WebUI,无需安装 Python、无需配置环境变量、无需额外下载浏览器内核)。
---
## 📁 项目目录结构
```text
FreePEP/
├── pep_core.py # 核心底层库(AES 解密、Playwright 爬取、PDF 合成)
├── webui.py # FastAPI WebUI 服务器与一体化前端界面
├── cli.py # 交互式与参数化 CLI 终端下载器
├── download_all.py # 按「学段/年级」层级全量下载脚本(支持断点续传与缓存清理)
├── build_exe.py # 一键打包发布 Windows EXE 独立便携包脚本
├── pep_crawler.py # 命令行测试与示例下载脚本
├── pep_catalog.json # 全量教材元数据本地缓存(自动生成)
├── requirements.txt # 项目 Python 依赖清单
├── README.md # 项目使用说明文档
├── temp_pages/ # 图片下载临时缓存目录(下载后自动清理/保留)
└── downloads/ # 生成的高清 PDF 默认存放目录
```
---
## 🔍 技术原理解析
1. 略。我是不明白为什么免费教材免费提供阅读不提供免费下载。
---
## ⚠️ 免责声明 (Disclaimer)
1. 本项目仅供 Python 爬虫技术交流、逆向工程学习与个人学习研究使用,严禁用于任何商业用途或盈利活动。
2. 本项目下载的所有教材版权均归**人民教育出版社(PEP)**及相关版权所有方所有。
3. 使用本项目时请控制请求频率,严禁进行任何可能对官方服务器造成过大负载的行为。请于下载后 24 小时内自行删除,如需长期使用请购买或支持官方正版出版物。
4. 使用者因违反版权或不当使用造成的一切法律纠纷与责任,均由使用者个人自行承担,与本项目作者无关。
---
## 📄 开源许可
本项目基于 [MIT License](LICENSE) 协议开源。欢迎提交 Issue 与 Pull Request!
+51
View File
@@ -0,0 +1,51 @@
# -*- mode: python ; coding: utf-8 -*-
from PyInstaller.utils.hooks import collect_all
datas = []
binaries = []
hiddenimports = ['uvicorn.logging', 'uvicorn.loops', 'uvicorn.loops.auto', 'uvicorn.protocols', 'uvicorn.protocols.http', 'uvicorn.protocols.http.auto', 'uvicorn.protocols.websockets', 'uvicorn.protocols.websockets.auto', 'uvicorn.lifespans', 'uvicorn.lifespans.on']
tmp_ret = collect_all('playwright')
datas += tmp_ret[0]; binaries += tmp_ret[1]; hiddenimports += tmp_ret[2]
a = Analysis(
['webui.py'],
pathex=[],
binaries=binaries,
datas=datas,
hiddenimports=hiddenimports,
hookspath=[],
hooksconfig={},
runtime_hooks=[],
excludes=[],
noarchive=False,
optimize=0,
)
pyz = PYZ(a.pure)
exe = EXE(
pyz,
a.scripts,
[],
exclude_binaries=True,
name='FreePEP',
debug=False,
bootloader_ignore_signals=False,
strip=False,
upx=True,
console=True,
disable_windowed_traceback=False,
argv_emulation=False,
target_arch=None,
codesign_identity=None,
entitlements_file=None,
)
coll = COLLECT(
exe,
a.binaries,
a.datas,
strip=False,
upx=True,
upx_exclude=[],
name='FreePEP',
)
+33 -2
View File
@@ -20,6 +20,26 @@
---
# 2026/09/19 FreePEP 1.4
### 🚀 功能新增与改进 (Features & Improvements)
1. **支持下载原图高清版本 (Large Resolution)**:
- 核心下载器支持抓取 PEP 官方 `large` 目录高清原图(分辨率提升至 2174×3071),并自带 404 自动回退机制;
- WebUI 新增“下载高清原图版本”快捷勾选项,CLI 新增 `--high-res / --hd` 命令行参数与交互选择。
2. **文档补充**:
- README 新增常见问题排查(FAQ),针对 macOS 环境提示补装 `playwright install chromium`。
### 🐛 Bug 修复与代码重构 (Bug Fixes & Refactoring)
1. **空值异常修复**:修复 `cli.py` 与 `webui.py` 中因元数据字段 `nj: null` 导致的 `AttributeError: 'NoneType'
object has no attribute 'strip'` 隐蔽崩溃。
2. **架构重构**:将 `download_all.py` 内部嵌套的 `is_match_xd` 函数提取为模块顶层函数,提升可复用性与可测试性。
### 🧪 自动化测试与 CI 护栏 (Testing & CI)
1. **分层单元测试**:新建 `tests/test_all.py`(32 个测试用例全部通过),涵盖纯函数映射、排序权重、学段过滤、AES-
128-CBC 加解密 Round-trip、下载器离线快路径与 FastAPI WebUI 接口测试。
2. **开发依赖分离**:新增 `requirements-dev.txt`(`pytest`、`httpx`),避免污染生产/打包依赖。
3. **持续集成配置**:新增 GitHub Actions CI 工作流 (`.github/workflows/ci.yml`),在 Python 3.9~3.12
矩阵下自动运行自动化测试。
## ✨ 核心特性
@@ -104,6 +124,7 @@ python cli.py --xd "小学(六三学制)" --xk "语文" --nj "一年级" -o
| `--xk` | 指定学科 | `语文`、`数学`、`英语`、`物理`、`化学`、`历史`、`道德与法治` 等 |
| `--nj` | 指定年级/册次 | `一年级`、`二年级` ... `九年级`、`专项`、`必修` 等 |
| `--search`, `-s` | 关键词全局搜索 | `必修`、`高一`、`地理` |
| `--high-res`, `--hd` | 下载高清原图版本 (`large`),默认普通版本 (`mobile`) | 无需参数 |
| `--clear-cache`, `-c` | 清理本地下载临时图片缓存 (`temp_pages`) | 无需参数 |
| `--refresh`, `-r` | 强制重新从官方服务器拉取解密最新目录 | 无需参数 |
| `--output`, `-o` | 指定 PDF 保存目录 | 默认: `./downloads` |
@@ -177,9 +198,19 @@ FreePEP/
---
## 🔍 技术原理解析
## ❓ 疑难解答 (FAQ)
1. 略。我是不明白为什么免费教材免费提供阅读不提供免费下载。
### Q: macOS 用户使用 WebUI 可以在线阅读,但点击下载后显示“任务完成”,而 `downloads` 目录没有任何 PDF 文件?
**原因分析**:
* **“在线看”正常**:点击“在线阅读”是由您的本机浏览器直接打开人教社公开阅读页面,不依赖后台 Python 浏览器驱动。
* **下载没有文件**:后台批量下载切片和过盾需要通过 Playwright 驱动无头 Chromium 浏览器。在 macOS 系统源码运行环境下,如果仅执行了 `pip install -r requirements.txt`,而**漏装了 Playwright 浏览器内核**,后台就会抛出 `Executable doesn't exist` 错误导致下载中断;而任务退出后界面可能误提示完成。
**解决方案**:
在 macOS 终端中激活当前 Python 环境,执行以下命令手动补全安装 Playwright 的 Chromium 内核:
```bash
playwright install chromium
```
> **排查提示**:如仍有问题,请查看运行 `python webui.py` 的终端控制台窗口,观察是否有详细的异常报错输出(如网络超时或环境异常)。
---
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+13 -5
View File
@@ -144,14 +144,17 @@ def interactive_mode():
print("[-] 未选择有效教材,退出。")
return
res_choice = input("\n请选择画质 [1] 普通清晰度 (默认) [2] 高清大图 (large): ").strip()
is_high_res = (res_choice == "2")
out_dir = os.path.abspath("./downloads")
print(f"\n🚀 即将开始下载 {len(to_download)} 本教材,基础保存目录: {out_dir}")
print(f"\n🚀 即将开始下载 {len(to_download)} 本教材 ({'高清版本' if is_high_res else '普通版本'}),基础保存目录: {out_dir}")
print(f"🌲 采用「学段 ➔ 年级」两层子目录分类保存 (已存在 PDF 自动跳过)")
downloader = PepDownloader(headless=True, output_dir=out_dir)
for idx, b in enumerate(to_download, 1):
xd = b.get("xd", "其他学段")
nj = b.get("nj", "通用").strip() or "通用"
nj = (b.get("nj") or "通用").strip() or "通用"
import re
safe_xd = re.sub(r'[\/:*?"<>|]', '_', xd).strip()
safe_nj = re.sub(r'[\/:*?"<>|]', '_', nj).strip()
@@ -165,7 +168,8 @@ def interactive_mode():
custom_title=b.get("title"),
sub_dir=sub_dir,
skip_if_exists=True,
clean_temp=True
clean_temp=True,
high_res=is_high_res
)
print("\n🎉 全部选定任务执行完毕!")
@@ -198,16 +202,18 @@ def cli_args_mode(args):
out_dir = os.path.abspath(args.output)
downloader = PepDownloader(headless=True, output_dir=out_dir)
use_tree = not args.flat
is_high_res = args.high_res
print(f"\n📂 保存根目录: {out_dir}")
print(f"🌲 目录结构: {'按「学段/年级」两层子目录' if use_tree else '全部平铺在根目录'}")
print(f"🖼️ 图像规格: {'高清大图模式 (large)' if is_high_res else '普通模式 (mobile)'}")
for idx, b in enumerate(matched, 1):
sub_dir = None
if use_tree:
import re
safe_xd = re.sub(r'[\/:*?"<>|]', '_', b.get("xd", "其他学段")).strip()
safe_nj = re.sub(r'[\/:*?"<>|]', '_', b.get("nj", "通用") or "通用").strip()
safe_nj = re.sub(r'[\/:*?"<>|]', '_', (b.get("nj") or "通用").strip() or "通用").strip()
sub_dir = os.path.join(safe_xd, safe_nj)
print(f"\n[{idx}/{len(matched)}] 正在下载: 《{b.get('title')}》...")
@@ -216,7 +222,8 @@ def cli_args_mode(args):
custom_title=b.get("title"),
sub_dir=sub_dir,
skip_if_exists=True,
clean_temp=True
clean_temp=True,
high_res=is_high_res
)
print(f"\n🎉 下载完成!文件已保存至: {out_dir}")
@@ -229,6 +236,7 @@ def main():
parser.add_argument("--xk", help="指定学科(如:语文、数学、英语、物理等)")
parser.add_argument("--nj", help="指定年级(如:一年级、七年级、必修等)")
parser.add_argument("--search", "-s", help="全局搜索关键词")
parser.add_argument("--high-res", "--hd", action="store_true", help="下载高清大图版本 (large),默认普通版本 (mobile)")
parser.add_argument("--output", "-o", default="./downloads", help="PDF 文件保存目录 (默认: ./downloads)")
parser.add_argument("--flat", action="store_true", help="平铺存放在根目录下(默认自动按「学段/年级」两层子目录分类)")
parser.add_argument("--yes", "-y", action="store_true", help="免确认直接开始下载")
+33 -26
View File
@@ -29,6 +29,7 @@ import time
import random
import argparse
import threading
from typing import List
from concurrent.futures import ThreadPoolExecutor, as_completed
from pep_core import PepCatalog, PepDownloader, normalize_xd, get_base_dir
@@ -41,27 +42,8 @@ def sanitize_filename(name: str) -> str:
return re.sub(r'[\/:*?"<>|]', '_', name).strip()
def run_download_all(output_dir: str = "./downloads",
target_xd: str = None,
target_nj: str = None,
max_workers: int = 3,
delay: float = 1.0,
clean_temp: bool = True):
print("=" * 70)
print(" 📚 人民教育出版社 (PEP) 全量教材多线程层级下载器")
print("=" * 70)
# 1. 获取全量教材目录数据
print("[*] 正在加载教材全量目录数据...")
all_books = PepCatalog.fetch_and_decrypt_all()
print(f"[+] 成功获取全量教材数据库,共 {len(all_books)} 本。")
# 2. 条件过滤(支持多个学段,逗号分隔,如:"义务教育(六三学制),义务教育(五四学制),高中")
target_xds = []
if target_xd and target_xd != "全部":
target_xds = [x.strip() for x in re.split(r'[,,|/]', target_xd) if x.strip()]
def is_match_xd(b_xd, b_xdtype):
def is_match_xd(b_xd: str, b_xdtype: str, target_xds: List[str]) -> bool:
"""判断教材学段是否与目标学段过滤条件匹配"""
if not target_xds:
return True
for t in target_xds:
@@ -76,13 +58,35 @@ def run_download_all(output_dir: str = "./downloads",
return True
return False
def run_download_all(output_dir: str = "./downloads",
target_xd: str = None,
target_nj: str = None,
max_workers: int = 3,
delay: float = 1.0,
clean_temp: bool = True,
high_res: bool = False):
print("=" * 70)
print(" 📚 人民教育出版社 (PEP) 全量教材多线程层级下载器")
print("=" * 70)
# 1. 获取全量教材目录数据
print("[*] 正在加载教材全量目录数据...")
all_books = PepCatalog.fetch_and_decrypt_all()
print(f"[+] 成功获取全量教材数据库,共 {len(all_books)} 本。")
# 2. 条件过滤(支持多个学段,逗号分隔,如:"义务教育(六三学制),义务教育(五四学制),高中")
target_xds = []
if target_xd and target_xd != "全部":
target_xds = [x.strip() for x in re.split(r'[,,|/]', target_xd) if x.strip()]
books_to_download = []
for b in all_books:
xd = normalize_xd(b.get("xd", "其他学段"))
xdtype = b.get("xdtype", "").strip()
nj = b.get("nj", "通用").strip() or "通用"
nj = (b.get("nj") or "通用").strip() or "通用"
if not is_match_xd(xd, xdtype):
if not is_match_xd(xd, xdtype, target_xds):
continue
if target_nj and target_nj != "全部" and nj != target_nj:
continue
@@ -130,7 +134,7 @@ def run_download_all(output_dir: str = "./downloads",
else:
stage_dir_name = xd
nj = b.get("nj", "通用").strip() or "通用"
nj = (b.get("nj") or "通用").strip() or "通用"
safe_xd = sanitize_filename(stage_dir_name)
safe_nj = sanitize_filename(nj)
@@ -161,7 +165,8 @@ def run_download_all(output_dir: str = "./downloads",
sub_dir=sub_dir,
skip_if_exists=True,
clean_temp=clean_temp,
quiet=True
quiet=True,
high_res=high_res
)
with lock:
@@ -215,6 +220,7 @@ def main():
parser.add_argument("--xd", help="只下载指定学段(如:小学(六三学制)、初中(六三学制)、高中 等)")
parser.add_argument("--nj", help="只下载指定年级(如:一年级、七年级 等)")
parser.add_argument("--delay", "-d", type=float, default=0.5, help="单线程任务间休眠秒数 (默认: 0.5 秒)")
parser.add_argument("--high-res", "--hd", action="store_true", help="下载高清大图版本 (large),默认普通版本 (mobile)")
parser.add_argument("--keep-temp", action="store_true", help="保留单页切片图片缓存(默认会自动删除以节约空间)")
args = parser.parse_args()
@@ -225,7 +231,8 @@ def main():
target_nj=args.nj,
max_workers=args.workers,
delay=args.delay,
clean_temp=not args.keep_temp
clean_temp=not args.keep_temp,
high_res=args.high_res
)
+33 -5
View File
@@ -391,7 +391,8 @@ class PepDownloader:
log_cb: Optional[Callable[[str], None]] = None,
skip_if_exists: bool = True,
clean_temp: bool = True,
quiet: bool = False) -> Optional[str]:
quiet: bool = False,
high_res: bool = False) -> Optional[str]:
"""
下载单本教材
:param book_id: 教材 ID(如 1284001101241)
@@ -402,6 +403,7 @@ class PepDownloader:
:param skip_if_exists: 若本地已存在完整 PDF 则自动跳过
:param clean_temp: 合成 PDF 后自动删除该书的切片图片以节约磁盘空间
:param quiet: 静默模式,不打印各页下载过程,仅在关键节点或报错时提示
:param high_res: 是否下载高清版本图片 (large),默认普通版本 (mobile)
:return: 生成的 PDF 绝对路径
"""
target_dir = os.path.join(self.output_dir, sub_dir) if sub_dir else self.output_dir
@@ -498,17 +500,20 @@ class PepDownloader:
browser.close()
return None
res_type = "large" if high_res else "mobile"
if not quiet:
info_msg = f"[+] 教材: 《{safe_title}》 | 总页数: {total_pages} 页"
mode_str = "高清模式 (large)" if high_res else "普通模式 (mobile)"
info_msg = f"[+] 教材: 《{safe_title}》 | 规格: {mode_str} | 总页数: {total_pages} 页"
if log_cb: log_cb(info_msg)
else: print(info_msg)
temp_dir = os.path.join(get_base_dir(), "temp_pages", book_id)
temp_dir = os.path.join(get_base_dir(), "temp_pages", f"{book_id}_{res_type}")
os.makedirs(temp_dir, exist_ok=True)
image_files = []
for page_num in range(1, total_pages + 1):
img_url = f"https://book.pep.com.cn/{book_id}/files/mobile/{page_num}.jpg"
img_url = f"https://book.pep.com.cn/{book_id}/files/{res_type}/{page_num}.jpg"
fallback_url = f"https://book.pep.com.cn/{book_id}/files/mobile/{page_num}.jpg" if high_res else None
img_path = os.path.join(temp_dir, f"{page_num}.jpg")
# 本地已有合法 JPEG 则跳过
@@ -521,6 +526,8 @@ class PepDownloader:
download_success = False
for retry in range(4):
# 尝试下载图片(如果为高清模式且返回404等错误,自动回退到普通 mobile 版本)
current_fetch_url = img_url
res = page.evaluate("""async (url) => {
try {
const resp = await fetch(url);
@@ -538,7 +545,28 @@ class PepDownloader:
} catch (e) {
return { status: 500, error: e.toString() };
}
}""", img_url)
}""", current_fetch_url)
# 如果高清版资源不存在(404),尝试回退到普通版
if fallback_url and res.get("status") == 404:
res = page.evaluate("""async (url) => {
try {
const resp = await fetch(url);
const ctype = resp.headers.get('content-type') || '';
const blob = await resp.blob();
return new Promise((resolve) => {
const reader = new FileReader();
reader.onloadend = () => resolve({
status: resp.status,
ctype: ctype,
data: reader.result
});
reader.readAsDataURL(blob);
});
} catch (e) {
return { status: 500, error: e.toString() };
}
}""", fallback_url)
data_uri = res.get("data", "")
if res.get("status") == 200 and data_uri and ("image" in res.get("ctype", "") or data_uri.startswith("data:image")):
+2
View File
@@ -0,0 +1,2 @@
pytest>=7.0.0
httpx>=0.23.0
+234
View File
@@ -0,0 +1,234 @@
"""
单元测试集合 (tests/test_all.py)
涵盖:
1. 纯函数与排序、映射规则测试 (表驱动)
2. 学段匹配逻辑与空值安全性测试 (is_match_xd, nj None-safety)
3. 本地教材目录检索与结构化分类测试 (PepCatalog filter & structure)
4. AES-128-CBC 加解密链路 round-trip 与正则匹配测试
5. PepDownloader skip_if_exists 离线快路径测试
6. WebUI FastAPI API 接口测试
"""
import os
import sys
import json
import base64
import binascii
import tempfile
import pytest
from cryptography.hazmat.primitives.ciphers import Cipher, algorithms, modes
from cryptography.hazmat.primitives import padding
from fastapi.testclient import TestClient
# 将项目根目录加入 sys.path
BASE_DIR = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
if BASE_DIR not in sys.path:
sys.path.insert(0, BASE_DIR)
from pep_core import (
normalize_xd,
sort_xd_key,
sort_xk_key,
sort_nj_key,
map_book_xd,
PepCatalog,
PepDownloader,
XD_ORDER,
XK_ORDER_PREFIX,
NJ_ORDER,
)
from download_all import sanitize_filename, is_match_xd
from webui import app
# ==========================================
# 1. 纯函数测试
# ==========================================
@pytest.mark.parametrize("raw, expected", [
("小学", "小学(六三学制)"),
("初中", "初中(六三学制)"),
("小学(五·四学制)", "小学(五四学制)"),
("小学(五四学制)", "小学(五四学制)"),
("初中(五·四学制)", "初中(五·四学制)"),
("初中(五四学制)", "初中(五·四学制)"),
("高中", "高中"),
("", "其他"),
(None, "其他"),
])
def test_normalize_xd(raw, expected):
assert normalize_xd(raw) == expected
def test_sort_keys():
# 学段排序测试
sorted_xds = sorted(XD_ORDER, key=sort_xd_key)
assert sorted_xds == XD_ORDER
# 未知学段排在后面
assert sort_xd_key("未知学段")[0] == 1
# 学科前缀优先排序
sorted_xks = sorted(XK_ORDER_PREFIX, key=sort_xk_key)
assert sorted_xks == XK_ORDER_PREFIX
assert sort_xk_key("未知学科")[0] == 1
# 年级排序
sorted_njs = sorted(NJ_ORDER, key=sort_nj_key)
assert sorted_njs == NJ_ORDER
@pytest.mark.parametrize("meta, expected", [
({"xd": "小学", "xdtype": "盲文"}, "盲校(盲文版)"),
({"xd": "初中", "xdtype": "低视力"}, "盲校(低视力版)"),
({"xd": "小学", "xdtype": "聋校"}, "聋校"),
({"xd": "小学", "xdtype": "培智"}, "培智学校"),
({"xd": "初中", "xdtype": "六三"}, "初中(六三学制)"),
({"xd": "小学", "xdtype": "六三"}, "小学(六三学制)"),
({"xd": "初中", "xdtype": "五四"}, "初中(五·四学制)"),
({"xd": "小学", "xdtype": "五四"}, "小学(五四学制)"),
({"xd": "高中", "xdtype": ""}, "高中"),
])
def test_map_book_xd(meta, expected):
assert map_book_xd(meta) == expected
def test_sanitize_filename():
assert sanitize_filename('test/book:name*1') == "test_book_name_1"
assert sanitize_filename('test/book:name*1?"<>|') == "test_book_name_1_____"
# ==========================================
# 2. 学段匹配与空值安全
# ==========================================
@pytest.mark.parametrize("b_xd, b_xdtype, targets, expected", [
("小学(六三学制)", "六三学制", ["义务教育(六三学制)"], True),
("初中(五·四学制)", "五四学制", ["义务教育(五四学制)"], True),
("高中", "普通高中", ["高中"], True),
("小学(六三学制)", "六三学制", ["高中"], False),
("小学", "六三", [], True), # 无过滤条件全部匹配
])
def test_is_match_xd(b_xd, b_xdtype, targets, expected):
assert is_match_xd(b_xd, b_xdtype, targets) == expected
def test_nj_null_safety():
"""测试 nj 为 None 时的安全性"""
b = {"id": "123", "title": "测试教材", "xd": "小学", "nj": None}
nj = (b.get("nj") or "通用").strip() or "通用"
assert nj == "通用"
# ==========================================
# 3. 目录检索与结构化分类 (读取本地 pep_catalog.json)
# ==========================================
def test_pep_catalog_filter_and_structure():
books = PepCatalog.fetch_and_decrypt_all()
assert len(books) > 0, "应成功从本地缓存读取教材数据"
# 测试条件过滤
filtered = PepCatalog.filter_books(xd="高中", xk="数学")
assert len(filtered) > 0
for b in filtered:
assert b["xd"] == "高中"
assert b["xk"] == "数学"
# 测试关键词搜索
kw_filtered = PepCatalog.filter_books(keyword="语文")
assert len(kw_filtered) > 0
# 测试结构体解析
structure = PepCatalog.get_structure()
assert "高中" in structure
assert "数学" in structure["高中"]["subjects"]
# ==========================================
# 4. AES-128-CBC 加解密链路 round-trip 测试
# ==========================================
def test_aes_round_trip():
key = PepCatalog.KEY
iv = PepCatalog.IV
sample_data = {"data": [{"id": "999999", "title": "单元测试教材"}]}
plain_bytes = json.dumps(sample_data).encode("utf-8")
# 模拟加密与 PKCS7 padding
padder = padding.PKCS7(128).padder()
padded_data = padder.update(plain_bytes) + padder.finalize()
cipher = Cipher(algorithms.AES(key), modes.CBC(iv))
encryptor = cipher.encryptor()
cipher_bytes = encryptor.update(padded_data) + encryptor.finalize()
hex_str = binascii.hexlify(cipher_bytes).decode("ascii").upper()
# 模拟从前端 JS 中匹配 hex_str
mock_js = f'var o, c = "{hex_str}";'.encode("ascii")
import re
m = re.search(rb'c\s*=\s*"([A-F0-9]+)"', mock_js)
assert m is not None
extracted_hex = m.group(1).decode("ascii")
extracted_cipher = binascii.unhexlify(extracted_hex)
# 执行解密
decryptor = cipher.decryptor()
decrypted_padded = decryptor.update(extracted_cipher) + decryptor.finalize()
unpadder = padding.PKCS7(128).unpadder()
decrypted_plain = unpadder.update(decrypted_padded) + unpadder.finalize()
res_json = json.loads(decrypted_plain.decode("utf-8"))
assert res_json == sample_data
# ==========================================
# 5. 下载器 skip_if_exists 离线快路径测试
# ==========================================
def test_downloader_skip_if_exists():
with tempfile.TemporaryDirectory() as tmp_dir:
downloader = PepDownloader(headless=True, output_dir=tmp_dir)
fake_pdf = os.path.join(tmp_dir, "测试教材.pdf")
# 写入大于 50KB 的假文件模拟已存在 PDF
with open(fake_pdf, "wb") as f:
f.write(b"%PDF-1.4 " + b"0" * 60000)
# 调用 download_book,由于文件已存在且 > 50KB,应直接秒退并返回路径,无需启动 Playwright 浏览器
result_path = downloader.download_book(
book_id="1384001301261",
custom_title="测试教材",
skip_if_exists=True
)
assert result_path == fake_pdf
# ==========================================
# 6. WebUI API 接口测试
# ==========================================
client = TestClient(app)
def test_webui_api_structure():
response = client.get("/api/structure")
assert response.status_code == 200
data = response.json()
assert isinstance(data, dict)
assert len(data) > 0
def test_webui_api_books():
response = client.post("/api/books", json={"xd": "高中", "xk": "数学", "nj": "全部", "keyword": ""})
assert response.status_code == 200
data = response.json()
assert "books" in data
assert data["total"] > 0
def test_webui_api_status():
response = client.get("/api/status")
assert response.status_code == 200
data = response.json()
assert "is_running" in data
assert "queue_len" in data
+25 -10
View File
@@ -58,16 +58,19 @@ class TaskManager:
"recent_logs": self.logs[-25:]
}
def add_tasks(self, books: List[Dict]):
def add_tasks(self, books: List[Dict], high_res: bool = False):
with self._lock:
existing_ids = {b["id"] for b in self.queue}
if self.current_book:
existing_ids.add(self.current_book["id"])
for b in books:
if b["id"] not in existing_ids:
self.queue.append(b)
item = dict(b)
item["_high_res"] = high_res
self.queue.append(item)
existing_ids.add(b["id"])
self.log(f"[*] 已添加 {len(books)} 本教材至下载队列。")
quality_str = "高清模式" if high_res else "普通模式"
self.log(f"[*] 已添加 {len(books)} 本教材至下载队列 ({quality_str})。")
def run_worker(self):
"""后台单线程顺序执行下载队列中的教材"""
@@ -104,14 +107,16 @@ class TaskManager:
self.log(txt)
try:
import re
xd = normalize_xd(book_to_download.get("xd", "其他学段"))
nj = book_to_download.get("nj", "通用").strip() or "通用"
nj = (book_to_download.get("nj") or "通用").strip() or "通用"
safe_xd = re.sub(r'[\/:*?"<>|]', '_', xd).strip()
safe_nj = re.sub(r'[\/:*?"<>|]', '_', nj).strip()
sub_dir = os.path.join(safe_xd, safe_nj)
is_high_res = book_to_download.get("_high_res", False)
self.log(f"==================================================")
self.log(f"[*] 开始下载教材: [{safe_xd}/{safe_nj}] 《{book_to_download.get('title')}》")
self.log(f"[*] 开始下载教材: [{safe_xd}/{safe_nj}] 《{book_to_download.get('title')}》 ({'高清版' if is_high_res else '普通版'})")
downloader.download_book(
book_id=book_to_download["id"],
custom_title=book_to_download.get("title"),
@@ -119,7 +124,8 @@ class TaskManager:
progress_cb=progress_callback,
log_cb=log_callback,
skip_if_exists=True,
clean_temp=True
clean_temp=True,
high_res=is_high_res
)
except Exception as e:
self.log(f"[-] 下载异常: {e}")
@@ -143,6 +149,7 @@ class FilterQuery(BaseModel):
class BatchDownloadRequest(BaseModel):
book_ids: List[str]
high_res: Optional[bool] = False
@app.get("/api/structure")
@@ -175,7 +182,7 @@ def add_download(req: BatchDownloadRequest):
all_books = {b["id"]: b for b in PepCatalog.fetch_and_decrypt_all()}
selected = [all_books[bid] for bid in req.book_ids if bid in all_books]
if selected:
task_manager.add_tasks(selected)
task_manager.add_tasks(selected, high_res=bool(req.high_res))
return {"status": "ok", "added_count": len(selected)}
@@ -324,7 +331,7 @@ def index_page():
<div class="space-y-4">
<!-- 批量操作条 -->
<div class="flex items-center justify-between bg-white border border-slate-200 px-4 py-3 rounded-lg text-xs">
<div class="flex flex-wrap items-center justify-between gap-3 bg-white border border-slate-200 px-4 py-3 rounded-lg text-xs">
<div class="flex items-center space-x-4">
<label class="flex items-center space-x-2 cursor-pointer select-none">
<input type="checkbox" id="selectAllCheckbox" onchange="toggleSelectAll()" class="rounded text-blue-600 focus:ring-blue-500">
@@ -334,11 +341,17 @@ def index_page():
<span class="text-slate-500">已选中 <b id="selectedCount" class="text-blue-600">0</b> 本</span>
</div>
<div class="flex items-center space-x-3">
<label class="flex items-center space-x-1.5 cursor-pointer select-none bg-slate-50 hover:bg-slate-100 border border-slate-200 px-2.5 py-1 rounded text-slate-700 font-medium transition" title="开启后将抓取 large 高清切片原图(若该教材无高清原图则自动降级为普通版)">
<input type="checkbox" id="highResToggle" class="rounded text-blue-600 focus:ring-blue-500">
<span>🌟 下载高清原图版本 (large)</span>
</label>
<button onclick="downloadSelected()" id="batchBtn" disabled
class="px-4 py-1.5 bg-blue-600 hover:bg-blue-700 disabled:bg-slate-300 disabled:cursor-not-allowed text-white font-medium rounded-md shadow-sm transition flex items-center space-x-1.5">
<span>📥 一键批量下载已选教材</span>
</button>
</div>
</div>
<!-- 教材卡片网格 -->
<div id="bookGrid" class="grid grid-cols-1 sm:grid-cols-2 md:grid-cols-3 lg:grid-cols-4 gap-4">
@@ -560,10 +573,11 @@ def index_page():
}
async function downloadSingle(id) {
const highRes = document.getElementById('highResToggle')?.checked || false;
await fetch('/api/download', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ book_ids: [id] })
body: JSON.stringify({ book_ids: [id], high_res: highRes })
});
pollStatus();
}
@@ -571,10 +585,11 @@ def index_page():
async function downloadSelected() {
const ids = Array.from(selectedBookIds);
if (ids.length === 0) return;
const highRes = document.getElementById('highResToggle')?.checked || false;
await fetch('/api/download', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ book_ids: ids })
body: JSON.stringify({ book_ids: ids, high_res: highRes })
});
selectedBookIds.clear();
updateSelectionUI();