diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml new file mode 100644 index 0000000..acecae0 --- /dev/null +++ b/.github/workflows/ci.yml @@ -0,0 +1,32 @@ +name: CI Tests + +on: + push: + branches: [ main, master ] + pull_request: + branches: [ main, master ] + +jobs: + test: + runs-on: ubuntu-latest + strategy: + matrix: + python-version: ["3.9", "3.10", "3.11", "3.12"] + + steps: + - uses: actions/checkout@v4 + + - name: Set up Python ${{ matrix.python-version }} + uses: actions/setup-python@v5 + with: + python-version: ${{ matrix.python-version }} + + - name: Install dependencies + run: | + python -m pip install --upgrade pip + pip install -r requirements.txt + pip install -r requirements-dev.txt + + - name: Run pytest + run: | + pytest -v tests/ diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..2247d5f --- /dev/null +++ b/.gitignore @@ -0,0 +1,2 @@ +/build +/dist diff --git a/FreePEP 📚 人教社中小学电子教材批量下载器.md b/FreePEP 📚 人教社中小学电子教材批量下载器.md deleted file mode 100644 index 909486b..0000000 --- a/FreePEP 📚 人教社中小学电子教材批量下载器.md +++ /dev/null @@ -1,209 +0,0 @@ -# FreePEP 📚 人教社中小学电子教材批量下载器 - -[](https://opensource.org/licenses/MIT) - -**FreePEP** 是一款专为[人民教育出版社中小学电子教材平台](https://jc.pep.com.cn/)开发的自动化教材解析、批量抓取与高清 PDF 合成工具。 - -提供**WebUI 界面**与**交互式命令行**,内置全量 780+ 本教材目录(数据截止到2026年08月31日)自动解密引擎与阿里云 WAF 滑块验证码自动破解机制,支持一键下载指定学段、学科、年级的全套教材并自动生成高清 PDF 文件。 - -**更新内容:** - -2026.09.04 FreePEP v1.1 命令行加入了批量下载功能,详情见**启动办法三** - - ---- - -## ✨ 核心特性 - -- 🎯 **双操作模式**: - - **现代化 WebUI 界面**:全响应式布局,还原官网层级筛选体验,支持复选框一键批量下载、实时进度条展示与本地保存目录快捷打开。 - - **交互式 CLI 终端**:支持数字菜单导航选择与命令行参数直达(脚本集成与服务器环境友好)。 - - 📦 **高清单页下载与 PDF 合成**:自动探测各教材实际总页数,批量下载高分辨率原始 JPG,并通过 Pillow 库自动合成为标准 PDF 文件。 - - ---- - -## 📸 界面预览 - -### WebUI 网页端 - - - ---- - -### CLI 网页端 - - - - -# 🚀 使用方式 - -## 🖥️ 方式一:(推荐,适合小白) - -直接下载发行包,解压缩后运行FreePEP.exe,在弹出的Web页面里面自行操作下载。 - -## 方式二:从源码启动 - -### 克隆仓库与安装依赖 - -```bash -# 克隆仓库 -git clone https://github.com/siknet/FreePEP.git -cd FreePEP - -# 安装 Python 依赖 -pip install -r requirements.txt - -# 安装 Playwright 所需的 Chromium 浏览器内核 -playwright install chromium -``` - -### 启动办法一:Webui - -在终端中执行以下命令,系统将自动启动本地服务并在默认浏览器中打开管理页面: - -```bash -python webui.py -``` - -- **访问地址**:`http://127.0.0.1:8000` -- **默认下载目录**:项目根目录下的 `./downloads` 文件夹(可在 Web 界面点击「📁 打开保存目录」直达)。 - ---- - -### 启动办法二:使用 CLI 命令行 - -#### 1. 交互式菜单模式 - -直接运行 `cli.py`,根据控制台提示逐步选择学段、学科与年级: - -```bash -python cli.py -``` - -#### 2. 参数直达模式(适合自动化与脚本调用) - -通过命令行参数直接指定条件过滤并自动开始下载: - -```bash -# 1. 下载小学一年级的所有学科教材 -python cli.py --xd "小学" --nj "一年级" -y - -# 2. 下载高中数学的所有必修/选修教材 -python cli.py --xd "高中" --xk "数学" -y - -# 3. 下载初中全部教材 -python cli.py --xd "初中" -y - -# 4. 全局关键词搜索并下载(如包含"物理"的所有教材) -python cli.py --search "物理" - -# 5. 清理本地临时图片缓存 -python cli.py --clear-cache - -# 6. 自定义 PDF 输出路径 -python cli.py --xd "小学(六三学制)" --xk "语文" --nj "一年级" -o "D:/Textbooks" -y -``` - -**参数说明**: - -| 参数 | 说明 | 示例 | -| :--- | :--- | :--- | -| `--xd` | 指定学段 | `小学(六三学制)`、`初中(六三学制)`、`小学(五四学制)`、`初中(五·四学制)`、`高中`、`培智学校`、`聋校`、`盲校(盲文版)`、`盲校(低视力版)` | -| `--xk` | 指定学科 | `语文`、`数学`、`英语`、`物理`、`化学`、`历史`、`道德与法治` 等 | -| `--nj` | 指定年级/册次 | `一年级`、`二年级` ... `九年级`、`专项`、`必修` 等 | -| `--search`, `-s` | 关键词全局搜索 | `必修`、`高一`、`地理` | -| `--clear-cache`, `-c` | 清理本地下载临时图片缓存 (`temp_pages`) | 无需参数 | -| `--refresh`, `-r` | 强制重新从官方服务器拉取解密最新目录 | 无需参数 | -| `--output`, `-o` | 指定 PDF 保存目录 | 默认: `./downloads` | -| `--yes`, `-y` | 跳过确认提示直接开始下载 | 开启免交互 | - -### 启动办法三:按「学段 ➔ 年级」层级批量多线程下载 (download_all.py) - -如果您希望将教材按照 **`学段/年级/`** 两层规范文件夹分类归档下载,直接运行专属的高速批量多线程下载脚本: - -```bash -# 1. 一键下载全网全量教材(默认 3 线程并发,按「学段/年级」两层子目录自动归档) -python download_all.py - -# 2. 启用多线程并发加速(例如 10 线程并行高速下载) -python download_all.py -w 10 - -# 3. 下载多个指定学段(支持逗号分隔,如:义务教育六三学制、五四学制与高中)并自定义保存路径 -python download_all.py --xd "义务教育(六三学制),义务教育(五四学制),高中" -w 10 -o "D:/人教社教材" - -# 4. 仅下载单个学段(如高中全部教材) -python download_all.py --xd "高中" - -# 5. 仅下载某个指定年级教材 -python download_all.py --nj "一年级" -``` - -**功能特性**: - -- 🚀 **多线程并发提速**:默认 **3 线程**并发同时下载,支持通过 `-w / --workers` 自定义并发线程数(如 5~10 线程),下载效率大幅提升。 -- 🎯 **支持多学段组合**:`--xd` 支持传入多个学段(逗号分隔),精准下载普通义务教育或高中学段,自动跳过特教教材。 -- 🌲 **两层层级目录**:自动分类保存为 `downloads/<学段>/<年级>/<教材名>.pdf`,告别成百上千文件堆在单目录。 -- ⚡ **智能断点续传**:已完整下载的教材**自动秒跳过**,中途随时中断无缝继续,绝不重复下载。 -- 🧹 **切片自动清理**:每本教材合成 PDF 后自动清理临时图片分片,极大保护硬盘空间。 -- 💬 **清爽紧凑输出**:不打印繁琐单页过程,仅在教材下载合并完成或跳过时输出进度与提示。 - ---- - -## 📦 一键打包为 Windows 独立 EXE 发行版 - -如果您想将本项目打包成 **脱离 Python 环境** 的独立绿色软件包发布给普通用户,只需执行工作区自带的一键打包脚本: - -```bash -python build_exe.py -``` - -### 打包脚本会自动完成以下操作: - -1. 自动检查并安装 `PyInstaller` 编译工具。 -2. 将 WebUI 及所有 Python 依赖打包为独立可执行文件 `FreePEP.exe`。 -3. **自动提取并内嵌绿色便携版 Chromium 浏览器内核**至 `browsers/` 目录。 -4. 自动在 `dist/` 目录下生成 `FreePEP-Windows-x64.zip` 发行压缩包。 - -> **分发给用户使用**:用户下载压缩包后解压,**双击 `FreePEP.exe` 即可直接使用**(会自动弹出系统默认浏览器打开 WebUI,无需安装 Python、无需配置环境变量、无需额外下载浏览器内核)。 - ---- - -## 📁 项目目录结构 - -```text -FreePEP/ -├── pep_core.py # 核心底层库(AES 解密、Playwright 爬取、PDF 合成) -├── webui.py # FastAPI WebUI 服务器与一体化前端界面 -├── cli.py # 交互式与参数化 CLI 终端下载器 -├── download_all.py # 按「学段/年级」层级全量下载脚本(支持断点续传与缓存清理) -├── build_exe.py # 一键打包发布 Windows EXE 独立便携包脚本 -├── pep_crawler.py # 命令行测试与示例下载脚本 -├── pep_catalog.json # 全量教材元数据本地缓存(自动生成) -├── requirements.txt # 项目 Python 依赖清单 -├── README.md # 项目使用说明文档 -├── temp_pages/ # 图片下载临时缓存目录(下载后自动清理/保留) -└── downloads/ # 生成的高清 PDF 默认存放目录 -``` - ---- - -## 🔍 技术原理解析 - -1. 略。我是不明白为什么免费教材免费提供阅读不提供免费下载。 - ---- - -## ⚠️ 免责声明 (Disclaimer) - -1. 本项目仅供 Python 爬虫技术交流、逆向工程学习与个人学习研究使用,严禁用于任何商业用途或盈利活动。 -2. 本项目下载的所有教材版权均归**人民教育出版社(PEP)**及相关版权所有方所有。 -3. 使用本项目时请控制请求频率,严禁进行任何可能对官方服务器造成过大负载的行为。请于下载后 24 小时内自行删除,如需长期使用请购买或支持官方正版出版物。 -4. 使用者因违反版权或不当使用造成的一切法律纠纷与责任,均由使用者个人自行承担,与本项目作者无关。 - ---- - -## 📄 开源许可 - -本项目基于 [MIT License](LICENSE) 协议开源。欢迎提交 Issue 与 Pull Request! - diff --git a/FreePEP.spec b/FreePEP.spec new file mode 100644 index 0000000..68819dc --- /dev/null +++ b/FreePEP.spec @@ -0,0 +1,51 @@ +# -*- mode: python ; coding: utf-8 -*- +from PyInstaller.utils.hooks import collect_all + +datas = [] +binaries = [] +hiddenimports = ['uvicorn.logging', 'uvicorn.loops', 'uvicorn.loops.auto', 'uvicorn.protocols', 'uvicorn.protocols.http', 'uvicorn.protocols.http.auto', 'uvicorn.protocols.websockets', 'uvicorn.protocols.websockets.auto', 'uvicorn.lifespans', 'uvicorn.lifespans.on'] +tmp_ret = collect_all('playwright') +datas += tmp_ret[0]; binaries += tmp_ret[1]; hiddenimports += tmp_ret[2] + + +a = Analysis( + ['webui.py'], + pathex=[], + binaries=binaries, + datas=datas, + hiddenimports=hiddenimports, + hookspath=[], + hooksconfig={}, + runtime_hooks=[], + excludes=[], + noarchive=False, + optimize=0, +) +pyz = PYZ(a.pure) + +exe = EXE( + pyz, + a.scripts, + [], + exclude_binaries=True, + name='FreePEP', + debug=False, + bootloader_ignore_signals=False, + strip=False, + upx=True, + console=True, + disable_windowed_traceback=False, + argv_emulation=False, + target_arch=None, + codesign_identity=None, + entitlements_file=None, +) +coll = COLLECT( + exe, + a.binaries, + a.datas, + strip=False, + upx=True, + upx_exclude=[], + name='FreePEP', +) diff --git a/README.md b/README.md index 428e31a..2959c72 100644 --- a/README.md +++ b/README.md @@ -20,7 +20,27 @@ --- +# 2026/09/19 FreePEP 1.4 +### 🚀 功能新增与改进 (Features & Improvements) + 1. **支持下载原图高清版本 (Large Resolution)**: + - 核心下载器支持抓取 PEP 官方 `large` 目录高清原图(分辨率提升至 2174×3071),并自带 404 自动回退机制; + - WebUI 新增“下载高清原图版本”快捷勾选项,CLI 新增 `--high-res / --hd` 命令行参数与交互选择。 + 2. **文档补充**: + - README 新增常见问题排查(FAQ),针对 macOS 环境提示补装 `playwright install chromium`。 + ### 🐛 Bug 修复与代码重构 (Bug Fixes & Refactoring) + 1. **空值异常修复**:修复 `cli.py` 与 `webui.py` 中因元数据字段 `nj: null` 导致的 `AttributeError: 'NoneType' + object has no attribute 'strip'` 隐蔽崩溃。 + 2. **架构重构**:将 `download_all.py` 内部嵌套的 `is_match_xd` 函数提取为模块顶层函数,提升可复用性与可测试性。 + + ### 🧪 自动化测试与 CI 护栏 (Testing & CI) + 1. **分层单元测试**:新建 `tests/test_all.py`(32 个测试用例全部通过),涵盖纯函数映射、排序权重、学段过滤、AES- + 128-CBC 加解密 Round-trip、下载器离线快路径与 FastAPI WebUI 接口测试。 + 2. **开发依赖分离**:新增 `requirements-dev.txt`(`pytest`、`httpx`),避免污染生产/打包依赖。 + 3. **持续集成配置**:新增 GitHub Actions CI 工作流 (`.github/workflows/ci.yml`),在 Python 3.9~3.12 + 矩阵下自动运行自动化测试。 + + ## ✨ 核心特性 - 🎯 **双操作模式**: @@ -104,6 +124,7 @@ python cli.py --xd "小学(六三学制)" --xk "语文" --nj "一年级" -o | `--xk` | 指定学科 | `语文`、`数学`、`英语`、`物理`、`化学`、`历史`、`道德与法治` 等 | | `--nj` | 指定年级/册次 | `一年级`、`二年级` ... `九年级`、`专项`、`必修` 等 | | `--search`, `-s` | 关键词全局搜索 | `必修`、`高一`、`地理` | +| `--high-res`, `--hd` | 下载高清原图版本 (`large`),默认普通版本 (`mobile`) | 无需参数 | | `--clear-cache`, `-c` | 清理本地下载临时图片缓存 (`temp_pages`) | 无需参数 | | `--refresh`, `-r` | 强制重新从官方服务器拉取解密最新目录 | 无需参数 | | `--output`, `-o` | 指定 PDF 保存目录 | 默认: `./downloads` | @@ -177,9 +198,19 @@ FreePEP/ --- -## 🔍 技术原理解析 +## ❓ 疑难解答 (FAQ) -1. 略。我是不明白为什么免费教材免费提供阅读不提供免费下载。 +### Q: macOS 用户使用 WebUI 可以在线阅读,但点击下载后显示“任务完成”,而 `downloads` 目录没有任何 PDF 文件? +**原因分析**: +* **“在线看”正常**:点击“在线阅读”是由您的本机浏览器直接打开人教社公开阅读页面,不依赖后台 Python 浏览器驱动。 +* **下载没有文件**:后台批量下载切片和过盾需要通过 Playwright 驱动无头 Chromium 浏览器。在 macOS 系统源码运行环境下,如果仅执行了 `pip install -r requirements.txt`,而**漏装了 Playwright 浏览器内核**,后台就会抛出 `Executable doesn't exist` 错误导致下载中断;而任务退出后界面可能误提示完成。 + +**解决方案**: +在 macOS 终端中激活当前 Python 环境,执行以下命令手动补全安装 Playwright 的 Chromium 内核: +```bash +playwright install chromium +``` +> **排查提示**:如仍有问题,请查看运行 `python webui.py` 的终端控制台窗口,观察是否有详细的异常报错输出(如网络超时或环境异常)。 --- diff --git a/__pycache__/cli.cpython-311.pyc b/__pycache__/cli.cpython-311.pyc new file mode 100644 index 0000000..b51e152 Binary files /dev/null and b/__pycache__/cli.cpython-311.pyc differ diff --git a/__pycache__/download_all.cpython-311.pyc b/__pycache__/download_all.cpython-311.pyc new file mode 100644 index 0000000..a23988b Binary files /dev/null and b/__pycache__/download_all.cpython-311.pyc differ diff --git a/__pycache__/pep_core.cpython-311.pyc b/__pycache__/pep_core.cpython-311.pyc new file mode 100644 index 0000000..7af0ef9 Binary files /dev/null and b/__pycache__/pep_core.cpython-311.pyc differ diff --git a/__pycache__/webui.cpython-311.pyc b/__pycache__/webui.cpython-311.pyc new file mode 100644 index 0000000..eef0098 Binary files /dev/null and b/__pycache__/webui.cpython-311.pyc differ diff --git a/cli.py b/cli.py index 0a981ed..dbb4074 100644 --- a/cli.py +++ b/cli.py @@ -144,14 +144,17 @@ def interactive_mode(): print("[-] 未选择有效教材,退出。") return + res_choice = input("\n请选择画质 [1] 普通清晰度 (默认) [2] 高清大图 (large): ").strip() + is_high_res = (res_choice == "2") + out_dir = os.path.abspath("./downloads") - print(f"\n🚀 即将开始下载 {len(to_download)} 本教材,基础保存目录: {out_dir}") + print(f"\n🚀 即将开始下载 {len(to_download)} 本教材 ({'高清版本' if is_high_res else '普通版本'}),基础保存目录: {out_dir}") print(f"🌲 采用「学段 ➔ 年级」两层子目录分类保存 (已存在 PDF 自动跳过)") downloader = PepDownloader(headless=True, output_dir=out_dir) for idx, b in enumerate(to_download, 1): xd = b.get("xd", "其他学段") - nj = b.get("nj", "通用").strip() or "通用" + nj = (b.get("nj") or "通用").strip() or "通用" import re safe_xd = re.sub(r'[\/:*?"<>|]', '_', xd).strip() safe_nj = re.sub(r'[\/:*?"<>|]', '_', nj).strip() @@ -165,7 +168,8 @@ def interactive_mode(): custom_title=b.get("title"), sub_dir=sub_dir, skip_if_exists=True, - clean_temp=True + clean_temp=True, + high_res=is_high_res ) print("\n🎉 全部选定任务执行完毕!") @@ -198,16 +202,18 @@ def cli_args_mode(args): out_dir = os.path.abspath(args.output) downloader = PepDownloader(headless=True, output_dir=out_dir) use_tree = not args.flat + is_high_res = args.high_res print(f"\n📂 保存根目录: {out_dir}") print(f"🌲 目录结构: {'按「学段/年级」两层子目录' if use_tree else '全部平铺在根目录'}") + print(f"🖼️ 图像规格: {'高清大图模式 (large)' if is_high_res else '普通模式 (mobile)'}") for idx, b in enumerate(matched, 1): sub_dir = None if use_tree: import re safe_xd = re.sub(r'[\/:*?"<>|]', '_', b.get("xd", "其他学段")).strip() - safe_nj = re.sub(r'[\/:*?"<>|]', '_', b.get("nj", "通用") or "通用").strip() + safe_nj = re.sub(r'[\/:*?"<>|]', '_', (b.get("nj") or "通用").strip() or "通用").strip() sub_dir = os.path.join(safe_xd, safe_nj) print(f"\n[{idx}/{len(matched)}] 正在下载: 《{b.get('title')}》...") @@ -216,7 +222,8 @@ def cli_args_mode(args): custom_title=b.get("title"), sub_dir=sub_dir, skip_if_exists=True, - clean_temp=True + clean_temp=True, + high_res=is_high_res ) print(f"\n🎉 下载完成!文件已保存至: {out_dir}") @@ -229,6 +236,7 @@ def main(): parser.add_argument("--xk", help="指定学科(如:语文、数学、英语、物理等)") parser.add_argument("--nj", help="指定年级(如:一年级、七年级、必修等)") parser.add_argument("--search", "-s", help="全局搜索关键词") + parser.add_argument("--high-res", "--hd", action="store_true", help="下载高清大图版本 (large),默认普通版本 (mobile)") parser.add_argument("--output", "-o", default="./downloads", help="PDF 文件保存目录 (默认: ./downloads)") parser.add_argument("--flat", action="store_true", help="平铺存放在根目录下(默认自动按「学段/年级」两层子目录分类)") parser.add_argument("--yes", "-y", action="store_true", help="免确认直接开始下载") diff --git a/download_all.py b/download_all.py index 02eab9e..1c8bdae 100644 --- a/download_all.py +++ b/download_all.py @@ -29,6 +29,7 @@ import time import random import argparse import threading +from typing import List from concurrent.futures import ThreadPoolExecutor, as_completed from pep_core import PepCatalog, PepDownloader, normalize_xd, get_base_dir @@ -41,12 +42,30 @@ def sanitize_filename(name: str) -> str: return re.sub(r'[\/:*?"<>|]', '_', name).strip() +def is_match_xd(b_xd: str, b_xdtype: str, target_xds: List[str]) -> bool: + """判断教材学段是否与目标学段过滤条件匹配""" + if not target_xds: + return True + for t in target_xds: + # 精确匹配或关键词包含匹配(如匹配 六三、五四、高中) + if t == b_xd or t == b_xdtype: + return True + if ("六三" in t and "六三" in b_xdtype) or ("六三" in t and "六三" in b_xd): + return True + if ("五四" in t or "五·四" in t) and ("五四" in b_xdtype or "五·四" in b_xdtype or "五四" in b_xd or "五·四" in b_xd): + return True + if ("高中" in t) and ("高中" in b_xd or "高中" in b_xdtype): + return True + return False + + def run_download_all(output_dir: str = "./downloads", target_xd: str = None, target_nj: str = None, max_workers: int = 3, delay: float = 1.0, - clean_temp: bool = True): + clean_temp: bool = True, + high_res: bool = False): print("=" * 70) print(" 📚 人民教育出版社 (PEP) 全量教材多线程层级下载器") print("=" * 70) @@ -61,28 +80,13 @@ def run_download_all(output_dir: str = "./downloads", if target_xd and target_xd != "全部": target_xds = [x.strip() for x in re.split(r'[,,|/]', target_xd) if x.strip()] - def is_match_xd(b_xd, b_xdtype): - if not target_xds: - return True - for t in target_xds: - # 精确匹配或关键词包含匹配(如匹配 六三、五四、高中) - if t == b_xd or t == b_xdtype: - return True - if ("六三" in t and "六三" in b_xdtype) or ("六三" in t and "六三" in b_xd): - return True - if ("五四" in t or "五·四" in t) and ("五四" in b_xdtype or "五·四" in b_xdtype or "五四" in b_xd or "五·四" in b_xd): - return True - if ("高中" in t) and ("高中" in b_xd or "高中" in b_xdtype): - return True - return False - books_to_download = [] for b in all_books: xd = normalize_xd(b.get("xd", "其他学段")) xdtype = b.get("xdtype", "").strip() - nj = b.get("nj", "通用").strip() or "通用" + nj = (b.get("nj") or "通用").strip() or "通用" - if not is_match_xd(xd, xdtype): + if not is_match_xd(xd, xdtype, target_xds): continue if target_nj and target_nj != "全部" and nj != target_nj: continue @@ -130,7 +134,7 @@ def run_download_all(output_dir: str = "./downloads", else: stage_dir_name = xd - nj = b.get("nj", "通用").strip() or "通用" + nj = (b.get("nj") or "通用").strip() or "通用" safe_xd = sanitize_filename(stage_dir_name) safe_nj = sanitize_filename(nj) @@ -161,7 +165,8 @@ def run_download_all(output_dir: str = "./downloads", sub_dir=sub_dir, skip_if_exists=True, clean_temp=clean_temp, - quiet=True + quiet=True, + high_res=high_res ) with lock: @@ -215,6 +220,7 @@ def main(): parser.add_argument("--xd", help="只下载指定学段(如:小学(六三学制)、初中(六三学制)、高中 等)") parser.add_argument("--nj", help="只下载指定年级(如:一年级、七年级 等)") parser.add_argument("--delay", "-d", type=float, default=0.5, help="单线程任务间休眠秒数 (默认: 0.5 秒)") + parser.add_argument("--high-res", "--hd", action="store_true", help="下载高清大图版本 (large),默认普通版本 (mobile)") parser.add_argument("--keep-temp", action="store_true", help="保留单页切片图片缓存(默认会自动删除以节约空间)") args = parser.parse_args() @@ -225,7 +231,8 @@ def main(): target_nj=args.nj, max_workers=args.workers, delay=args.delay, - clean_temp=not args.keep_temp + clean_temp=not args.keep_temp, + high_res=args.high_res ) diff --git a/pep_core.py b/pep_core.py index 050e9d4..ea5caa5 100644 --- a/pep_core.py +++ b/pep_core.py @@ -391,7 +391,8 @@ class PepDownloader: log_cb: Optional[Callable[[str], None]] = None, skip_if_exists: bool = True, clean_temp: bool = True, - quiet: bool = False) -> Optional[str]: + quiet: bool = False, + high_res: bool = False) -> Optional[str]: """ 下载单本教材 :param book_id: 教材 ID(如 1284001101241) @@ -402,6 +403,7 @@ class PepDownloader: :param skip_if_exists: 若本地已存在完整 PDF 则自动跳过 :param clean_temp: 合成 PDF 后自动删除该书的切片图片以节约磁盘空间 :param quiet: 静默模式,不打印各页下载过程,仅在关键节点或报错时提示 + :param high_res: 是否下载高清版本图片 (large),默认普通版本 (mobile) :return: 生成的 PDF 绝对路径 """ target_dir = os.path.join(self.output_dir, sub_dir) if sub_dir else self.output_dir @@ -498,17 +500,20 @@ class PepDownloader: browser.close() return None + res_type = "large" if high_res else "mobile" if not quiet: - info_msg = f"[+] 教材: 《{safe_title}》 | 总页数: {total_pages} 页" + mode_str = "高清模式 (large)" if high_res else "普通模式 (mobile)" + info_msg = f"[+] 教材: 《{safe_title}》 | 规格: {mode_str} | 总页数: {total_pages} 页" if log_cb: log_cb(info_msg) else: print(info_msg) - temp_dir = os.path.join(get_base_dir(), "temp_pages", book_id) + temp_dir = os.path.join(get_base_dir(), "temp_pages", f"{book_id}_{res_type}") os.makedirs(temp_dir, exist_ok=True) image_files = [] for page_num in range(1, total_pages + 1): - img_url = f"https://book.pep.com.cn/{book_id}/files/mobile/{page_num}.jpg" + img_url = f"https://book.pep.com.cn/{book_id}/files/{res_type}/{page_num}.jpg" + fallback_url = f"https://book.pep.com.cn/{book_id}/files/mobile/{page_num}.jpg" if high_res else None img_path = os.path.join(temp_dir, f"{page_num}.jpg") # 本地已有合法 JPEG 则跳过 @@ -521,6 +526,8 @@ class PepDownloader: download_success = False for retry in range(4): + # 尝试下载图片(如果为高清模式且返回404等错误,自动回退到普通 mobile 版本) + current_fetch_url = img_url res = page.evaluate("""async (url) => { try { const resp = await fetch(url); @@ -538,7 +545,28 @@ class PepDownloader: } catch (e) { return { status: 500, error: e.toString() }; } - }""", img_url) + }""", current_fetch_url) + + # 如果高清版资源不存在(404),尝试回退到普通版 + if fallback_url and res.get("status") == 404: + res = page.evaluate("""async (url) => { + try { + const resp = await fetch(url); + const ctype = resp.headers.get('content-type') || ''; + const blob = await resp.blob(); + return new Promise((resolve) => { + const reader = new FileReader(); + reader.onloadend = () => resolve({ + status: resp.status, + ctype: ctype, + data: reader.result + }); + reader.readAsDataURL(blob); + }); + } catch (e) { + return { status: 500, error: e.toString() }; + } + }""", fallback_url) data_uri = res.get("data", "") if res.get("status") == 200 and data_uri and ("image" in res.get("ctype", "") or data_uri.startswith("data:image")): diff --git a/requirements-dev.txt b/requirements-dev.txt new file mode 100644 index 0000000..5c79490 --- /dev/null +++ b/requirements-dev.txt @@ -0,0 +1,2 @@ +pytest>=7.0.0 +httpx>=0.23.0 diff --git a/tests/__pycache__/test_all.cpython-311-pytest-9.1.1.pyc b/tests/__pycache__/test_all.cpython-311-pytest-9.1.1.pyc new file mode 100644 index 0000000..52f87c3 Binary files /dev/null and b/tests/__pycache__/test_all.cpython-311-pytest-9.1.1.pyc differ diff --git a/tests/__pycache__/test_all.cpython-39-pytest-8.4.2.pyc b/tests/__pycache__/test_all.cpython-39-pytest-8.4.2.pyc new file mode 100644 index 0000000..a87470a Binary files /dev/null and b/tests/__pycache__/test_all.cpython-39-pytest-8.4.2.pyc differ diff --git a/tests/test_all.py b/tests/test_all.py new file mode 100644 index 0000000..8164b35 --- /dev/null +++ b/tests/test_all.py @@ -0,0 +1,234 @@ +""" +单元测试集合 (tests/test_all.py) +涵盖: +1. 纯函数与排序、映射规则测试 (表驱动) +2. 学段匹配逻辑与空值安全性测试 (is_match_xd, nj None-safety) +3. 本地教材目录检索与结构化分类测试 (PepCatalog filter & structure) +4. AES-128-CBC 加解密链路 round-trip 与正则匹配测试 +5. PepDownloader skip_if_exists 离线快路径测试 +6. WebUI FastAPI API 接口测试 +""" + +import os +import sys +import json +import base64 +import binascii +import tempfile +import pytest +from cryptography.hazmat.primitives.ciphers import Cipher, algorithms, modes +from cryptography.hazmat.primitives import padding +from fastapi.testclient import TestClient + +# 将项目根目录加入 sys.path +BASE_DIR = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +if BASE_DIR not in sys.path: + sys.path.insert(0, BASE_DIR) + +from pep_core import ( + normalize_xd, + sort_xd_key, + sort_xk_key, + sort_nj_key, + map_book_xd, + PepCatalog, + PepDownloader, + XD_ORDER, + XK_ORDER_PREFIX, + NJ_ORDER, +) +from download_all import sanitize_filename, is_match_xd +from webui import app + + +# ========================================== +# 1. 纯函数测试 +# ========================================== + +@pytest.mark.parametrize("raw, expected", [ + ("小学", "小学(六三学制)"), + ("初中", "初中(六三学制)"), + ("小学(五·四学制)", "小学(五四学制)"), + ("小学(五四学制)", "小学(五四学制)"), + ("初中(五·四学制)", "初中(五·四学制)"), + ("初中(五四学制)", "初中(五·四学制)"), + ("高中", "高中"), + ("", "其他"), + (None, "其他"), +]) +def test_normalize_xd(raw, expected): + assert normalize_xd(raw) == expected + + +def test_sort_keys(): + # 学段排序测试 + sorted_xds = sorted(XD_ORDER, key=sort_xd_key) + assert sorted_xds == XD_ORDER + # 未知学段排在后面 + assert sort_xd_key("未知学段")[0] == 1 + + # 学科前缀优先排序 + sorted_xks = sorted(XK_ORDER_PREFIX, key=sort_xk_key) + assert sorted_xks == XK_ORDER_PREFIX + assert sort_xk_key("未知学科")[0] == 1 + + # 年级排序 + sorted_njs = sorted(NJ_ORDER, key=sort_nj_key) + assert sorted_njs == NJ_ORDER + + +@pytest.mark.parametrize("meta, expected", [ + ({"xd": "小学", "xdtype": "盲文"}, "盲校(盲文版)"), + ({"xd": "初中", "xdtype": "低视力"}, "盲校(低视力版)"), + ({"xd": "小学", "xdtype": "聋校"}, "聋校"), + ({"xd": "小学", "xdtype": "培智"}, "培智学校"), + ({"xd": "初中", "xdtype": "六三"}, "初中(六三学制)"), + ({"xd": "小学", "xdtype": "六三"}, "小学(六三学制)"), + ({"xd": "初中", "xdtype": "五四"}, "初中(五·四学制)"), + ({"xd": "小学", "xdtype": "五四"}, "小学(五四学制)"), + ({"xd": "高中", "xdtype": ""}, "高中"), +]) +def test_map_book_xd(meta, expected): + assert map_book_xd(meta) == expected + + +def test_sanitize_filename(): + assert sanitize_filename('test/book:name*1') == "test_book_name_1" + assert sanitize_filename('test/book:name*1?"<>|') == "test_book_name_1_____" + + +# ========================================== +# 2. 学段匹配与空值安全 +# ========================================== + +@pytest.mark.parametrize("b_xd, b_xdtype, targets, expected", [ + ("小学(六三学制)", "六三学制", ["义务教育(六三学制)"], True), + ("初中(五·四学制)", "五四学制", ["义务教育(五四学制)"], True), + ("高中", "普通高中", ["高中"], True), + ("小学(六三学制)", "六三学制", ["高中"], False), + ("小学", "六三", [], True), # 无过滤条件全部匹配 +]) +def test_is_match_xd(b_xd, b_xdtype, targets, expected): + assert is_match_xd(b_xd, b_xdtype, targets) == expected + + +def test_nj_null_safety(): + """测试 nj 为 None 时的安全性""" + b = {"id": "123", "title": "测试教材", "xd": "小学", "nj": None} + nj = (b.get("nj") or "通用").strip() or "通用" + assert nj == "通用" + + +# ========================================== +# 3. 目录检索与结构化分类 (读取本地 pep_catalog.json) +# ========================================== + +def test_pep_catalog_filter_and_structure(): + books = PepCatalog.fetch_and_decrypt_all() + assert len(books) > 0, "应成功从本地缓存读取教材数据" + + # 测试条件过滤 + filtered = PepCatalog.filter_books(xd="高中", xk="数学") + assert len(filtered) > 0 + for b in filtered: + assert b["xd"] == "高中" + assert b["xk"] == "数学" + + # 测试关键词搜索 + kw_filtered = PepCatalog.filter_books(keyword="语文") + assert len(kw_filtered) > 0 + + # 测试结构体解析 + structure = PepCatalog.get_structure() + assert "高中" in structure + assert "数学" in structure["高中"]["subjects"] + + +# ========================================== +# 4. AES-128-CBC 加解密链路 round-trip 测试 +# ========================================== + +def test_aes_round_trip(): + key = PepCatalog.KEY + iv = PepCatalog.IV + + sample_data = {"data": [{"id": "999999", "title": "单元测试教材"}]} + plain_bytes = json.dumps(sample_data).encode("utf-8") + + # 模拟加密与 PKCS7 padding + padder = padding.PKCS7(128).padder() + padded_data = padder.update(plain_bytes) + padder.finalize() + + cipher = Cipher(algorithms.AES(key), modes.CBC(iv)) + encryptor = cipher.encryptor() + cipher_bytes = encryptor.update(padded_data) + encryptor.finalize() + hex_str = binascii.hexlify(cipher_bytes).decode("ascii").upper() + + # 模拟从前端 JS 中匹配 hex_str + mock_js = f'var o, c = "{hex_str}";'.encode("ascii") + import re + m = re.search(rb'c\s*=\s*"([A-F0-9]+)"', mock_js) + assert m is not None + + extracted_hex = m.group(1).decode("ascii") + extracted_cipher = binascii.unhexlify(extracted_hex) + + # 执行解密 + decryptor = cipher.decryptor() + decrypted_padded = decryptor.update(extracted_cipher) + decryptor.finalize() + unpadder = padding.PKCS7(128).unpadder() + decrypted_plain = unpadder.update(decrypted_padded) + unpadder.finalize() + + res_json = json.loads(decrypted_plain.decode("utf-8")) + assert res_json == sample_data + + +# ========================================== +# 5. 下载器 skip_if_exists 离线快路径测试 +# ========================================== + +def test_downloader_skip_if_exists(): + with tempfile.TemporaryDirectory() as tmp_dir: + downloader = PepDownloader(headless=True, output_dir=tmp_dir) + fake_pdf = os.path.join(tmp_dir, "测试教材.pdf") + # 写入大于 50KB 的假文件模拟已存在 PDF + with open(fake_pdf, "wb") as f: + f.write(b"%PDF-1.4 " + b"0" * 60000) + + # 调用 download_book,由于文件已存在且 > 50KB,应直接秒退并返回路径,无需启动 Playwright 浏览器 + result_path = downloader.download_book( + book_id="1384001301261", + custom_title="测试教材", + skip_if_exists=True + ) + assert result_path == fake_pdf + + +# ========================================== +# 6. WebUI API 接口测试 +# ========================================== + +client = TestClient(app) + +def test_webui_api_structure(): + response = client.get("/api/structure") + assert response.status_code == 200 + data = response.json() + assert isinstance(data, dict) + assert len(data) > 0 + + +def test_webui_api_books(): + response = client.post("/api/books", json={"xd": "高中", "xk": "数学", "nj": "全部", "keyword": ""}) + assert response.status_code == 200 + data = response.json() + assert "books" in data + assert data["total"] > 0 + + +def test_webui_api_status(): + response = client.get("/api/status") + assert response.status_code == 200 + data = response.json() + assert "is_running" in data + assert "queue_len" in data diff --git a/webui.py b/webui.py index 9b278de..49941fb 100644 --- a/webui.py +++ b/webui.py @@ -58,16 +58,19 @@ class TaskManager: "recent_logs": self.logs[-25:] } - def add_tasks(self, books: List[Dict]): + def add_tasks(self, books: List[Dict], high_res: bool = False): with self._lock: existing_ids = {b["id"] for b in self.queue} if self.current_book: existing_ids.add(self.current_book["id"]) for b in books: if b["id"] not in existing_ids: - self.queue.append(b) + item = dict(b) + item["_high_res"] = high_res + self.queue.append(item) existing_ids.add(b["id"]) - self.log(f"[*] 已添加 {len(books)} 本教材至下载队列。") + quality_str = "高清模式" if high_res else "普通模式" + self.log(f"[*] 已添加 {len(books)} 本教材至下载队列 ({quality_str})。") def run_worker(self): """后台单线程顺序执行下载队列中的教材""" @@ -104,14 +107,16 @@ class TaskManager: self.log(txt) try: + import re xd = normalize_xd(book_to_download.get("xd", "其他学段")) - nj = book_to_download.get("nj", "通用").strip() or "通用" + nj = (book_to_download.get("nj") or "通用").strip() or "通用" safe_xd = re.sub(r'[\/:*?"<>|]', '_', xd).strip() safe_nj = re.sub(r'[\/:*?"<>|]', '_', nj).strip() sub_dir = os.path.join(safe_xd, safe_nj) + is_high_res = book_to_download.get("_high_res", False) self.log(f"==================================================") - self.log(f"[*] 开始下载教材: [{safe_xd}/{safe_nj}] 《{book_to_download.get('title')}》") + self.log(f"[*] 开始下载教材: [{safe_xd}/{safe_nj}] 《{book_to_download.get('title')}》 ({'高清版' if is_high_res else '普通版'})") downloader.download_book( book_id=book_to_download["id"], custom_title=book_to_download.get("title"), @@ -119,7 +124,8 @@ class TaskManager: progress_cb=progress_callback, log_cb=log_callback, skip_if_exists=True, - clean_temp=True + clean_temp=True, + high_res=is_high_res ) except Exception as e: self.log(f"[-] 下载异常: {e}") @@ -143,6 +149,7 @@ class FilterQuery(BaseModel): class BatchDownloadRequest(BaseModel): book_ids: List[str] + high_res: Optional[bool] = False @app.get("/api/structure") @@ -175,7 +182,7 @@ def add_download(req: BatchDownloadRequest): all_books = {b["id"]: b for b in PepCatalog.fetch_and_decrypt_all()} selected = [all_books[bid] for bid in req.book_ids if bid in all_books] if selected: - task_manager.add_tasks(selected) + task_manager.add_tasks(selected, high_res=bool(req.high_res)) return {"status": "ok", "added_count": len(selected)} @@ -324,7 +331,7 @@ def index_page():