v1.1 加入了命令行下的批量下载功能
现在可以一次把人教社的身体掏空了,实测最大10线程没问题。
This commit is contained in:
@@ -0,0 +1,209 @@
|
|||||||
|
# FreePEP 📚 人教社中小学电子教材批量下载器
|
||||||
|
|
||||||
|
[](https://opensource.org/licenses/MIT)
|
||||||
|
|
||||||
|
**FreePEP** 是一款专为[人民教育出版社中小学电子教材平台](https://jc.pep.com.cn/)开发的自动化教材解析、批量抓取与高清 PDF 合成工具。
|
||||||
|
|
||||||
|
提供**WebUI 界面**与**交互式命令行**,内置全量 780+ 本教材目录(数据截止到2026年08月31日)自动解密引擎与阿里云 WAF 滑块验证码自动破解机制,支持一键下载指定学段、学科、年级的全套教材并自动生成高清 PDF 文件。
|
||||||
|
|
||||||
|
**更新内容:**
|
||||||
|
|
||||||
|
2026.09.04 FreePEP v1.1 命令行加入了批量下载功能,详情见**启动办法三**
|
||||||
|
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## ✨ 核心特性
|
||||||
|
|
||||||
|
- 🎯 **双操作模式**:
|
||||||
|
- **现代化 WebUI 界面**:全响应式布局,还原官网层级筛选体验,支持复选框一键批量下载、实时进度条展示与本地保存目录快捷打开。
|
||||||
|
- **交互式 CLI 终端**:支持数字菜单导航选择与命令行参数直达(脚本集成与服务器环境友好)。
|
||||||
|
- 📦 **高清单页下载与 PDF 合成**:自动探测各教材实际总页数,批量下载高分辨率原始 JPG,并通过 Pillow 库自动合成为标准 PDF 文件。
|
||||||
|
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 📸 界面预览
|
||||||
|
|
||||||
|
### WebUI 网页端
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### CLI 网页端
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
|
||||||
|
# 🚀 使用方式
|
||||||
|
|
||||||
|
## 🖥️ 方式一:(推荐,适合小白)
|
||||||
|
|
||||||
|
直接下载发行包,解压缩后运行FreePEP.exe,在弹出的Web页面里面自行操作下载。
|
||||||
|
|
||||||
|
## 方式二:从源码启动
|
||||||
|
|
||||||
|
### 克隆仓库与安装依赖
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 克隆仓库
|
||||||
|
git clone https://github.com/siknet/FreePEP.git
|
||||||
|
cd FreePEP
|
||||||
|
|
||||||
|
# 安装 Python 依赖
|
||||||
|
pip install -r requirements.txt
|
||||||
|
|
||||||
|
# 安装 Playwright 所需的 Chromium 浏览器内核
|
||||||
|
playwright install chromium
|
||||||
|
```
|
||||||
|
|
||||||
|
### 启动办法一:Webui
|
||||||
|
|
||||||
|
在终端中执行以下命令,系统将自动启动本地服务并在默认浏览器中打开管理页面:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python webui.py
|
||||||
|
```
|
||||||
|
|
||||||
|
- **访问地址**:`http://127.0.0.1:8000`
|
||||||
|
- **默认下载目录**:项目根目录下的 `./downloads` 文件夹(可在 Web 界面点击「📁 打开保存目录」直达)。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 启动办法二:使用 CLI 命令行
|
||||||
|
|
||||||
|
#### 1. 交互式菜单模式
|
||||||
|
|
||||||
|
直接运行 `cli.py`,根据控制台提示逐步选择学段、学科与年级:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python cli.py
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 2. 参数直达模式(适合自动化与脚本调用)
|
||||||
|
|
||||||
|
通过命令行参数直接指定条件过滤并自动开始下载:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. 下载小学一年级的所有学科教材
|
||||||
|
python cli.py --xd "小学" --nj "一年级" -y
|
||||||
|
|
||||||
|
# 2. 下载高中数学的所有必修/选修教材
|
||||||
|
python cli.py --xd "高中" --xk "数学" -y
|
||||||
|
|
||||||
|
# 3. 下载初中全部教材
|
||||||
|
python cli.py --xd "初中" -y
|
||||||
|
|
||||||
|
# 4. 全局关键词搜索并下载(如包含"物理"的所有教材)
|
||||||
|
python cli.py --search "物理"
|
||||||
|
|
||||||
|
# 5. 清理本地临时图片缓存
|
||||||
|
python cli.py --clear-cache
|
||||||
|
|
||||||
|
# 6. 自定义 PDF 输出路径
|
||||||
|
python cli.py --xd "小学(六三学制)" --xk "语文" --nj "一年级" -o "D:/Textbooks" -y
|
||||||
|
```
|
||||||
|
|
||||||
|
**参数说明**:
|
||||||
|
|
||||||
|
| 参数 | 说明 | 示例 |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `--xd` | 指定学段 | `小学(六三学制)`、`初中(六三学制)`、`小学(五四学制)`、`初中(五·四学制)`、`高中`、`培智学校`、`聋校`、`盲校(盲文版)`、`盲校(低视力版)` |
|
||||||
|
| `--xk` | 指定学科 | `语文`、`数学`、`英语`、`物理`、`化学`、`历史`、`道德与法治` 等 |
|
||||||
|
| `--nj` | 指定年级/册次 | `一年级`、`二年级` ... `九年级`、`专项`、`必修` 等 |
|
||||||
|
| `--search`, `-s` | 关键词全局搜索 | `必修`、`高一`、`地理` |
|
||||||
|
| `--clear-cache`, `-c` | 清理本地下载临时图片缓存 (`temp_pages`) | 无需参数 |
|
||||||
|
| `--refresh`, `-r` | 强制重新从官方服务器拉取解密最新目录 | 无需参数 |
|
||||||
|
| `--output`, `-o` | 指定 PDF 保存目录 | 默认: `./downloads` |
|
||||||
|
| `--yes`, `-y` | 跳过确认提示直接开始下载 | 开启免交互 |
|
||||||
|
|
||||||
|
### 启动办法三:按「学段 ➔ 年级」层级批量多线程下载 (download_all.py)
|
||||||
|
|
||||||
|
如果您希望将教材按照 **`学段/年级/`** 两层规范文件夹分类归档下载,直接运行专属的高速批量多线程下载脚本:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. 一键下载全网全量教材(默认 3 线程并发,按「学段/年级」两层子目录自动归档)
|
||||||
|
python download_all.py
|
||||||
|
|
||||||
|
# 2. 启用多线程并发加速(例如 10 线程并行高速下载)
|
||||||
|
python download_all.py -w 10
|
||||||
|
|
||||||
|
# 3. 下载多个指定学段(支持逗号分隔,如:义务教育六三学制、五四学制与高中)并自定义保存路径
|
||||||
|
python download_all.py --xd "义务教育(六三学制),义务教育(五四学制),高中" -w 10 -o "D:/人教社教材"
|
||||||
|
|
||||||
|
# 4. 仅下载单个学段(如高中全部教材)
|
||||||
|
python download_all.py --xd "高中"
|
||||||
|
|
||||||
|
# 5. 仅下载某个指定年级教材
|
||||||
|
python download_all.py --nj "一年级"
|
||||||
|
```
|
||||||
|
|
||||||
|
**功能特性**:
|
||||||
|
|
||||||
|
- 🚀 **多线程并发提速**:默认 **3 线程**并发同时下载,支持通过 `-w / --workers` 自定义并发线程数(如 5~10 线程),下载效率大幅提升。
|
||||||
|
- 🎯 **支持多学段组合**:`--xd` 支持传入多个学段(逗号分隔),精准下载普通义务教育或高中学段,自动跳过特教教材。
|
||||||
|
- 🌲 **两层层级目录**:自动分类保存为 `downloads/<学段>/<年级>/<教材名>.pdf`,告别成百上千文件堆在单目录。
|
||||||
|
- ⚡ **智能断点续传**:已完整下载的教材**自动秒跳过**,中途随时中断无缝继续,绝不重复下载。
|
||||||
|
- 🧹 **切片自动清理**:每本教材合成 PDF 后自动清理临时图片分片,极大保护硬盘空间。
|
||||||
|
- 💬 **清爽紧凑输出**:不打印繁琐单页过程,仅在教材下载合并完成或跳过时输出进度与提示。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 📦 一键打包为 Windows 独立 EXE 发行版
|
||||||
|
|
||||||
|
如果您想将本项目打包成 **脱离 Python 环境** 的独立绿色软件包发布给普通用户,只需执行工作区自带的一键打包脚本:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python build_exe.py
|
||||||
|
```
|
||||||
|
|
||||||
|
### 打包脚本会自动完成以下操作:
|
||||||
|
|
||||||
|
1. 自动检查并安装 `PyInstaller` 编译工具。
|
||||||
|
2. 将 WebUI 及所有 Python 依赖打包为独立可执行文件 `FreePEP.exe`。
|
||||||
|
3. **自动提取并内嵌绿色便携版 Chromium 浏览器内核**至 `browsers/` 目录。
|
||||||
|
4. 自动在 `dist/` 目录下生成 `FreePEP-Windows-x64.zip` 发行压缩包。
|
||||||
|
|
||||||
|
> **分发给用户使用**:用户下载压缩包后解压,**双击 `FreePEP.exe` 即可直接使用**(会自动弹出系统默认浏览器打开 WebUI,无需安装 Python、无需配置环境变量、无需额外下载浏览器内核)。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 📁 项目目录结构
|
||||||
|
|
||||||
|
```text
|
||||||
|
FreePEP/
|
||||||
|
├── pep_core.py # 核心底层库(AES 解密、Playwright 爬取、PDF 合成)
|
||||||
|
├── webui.py # FastAPI WebUI 服务器与一体化前端界面
|
||||||
|
├── cli.py # 交互式与参数化 CLI 终端下载器
|
||||||
|
├── download_all.py # 按「学段/年级」层级全量下载脚本(支持断点续传与缓存清理)
|
||||||
|
├── build_exe.py # 一键打包发布 Windows EXE 独立便携包脚本
|
||||||
|
├── pep_crawler.py # 命令行测试与示例下载脚本
|
||||||
|
├── pep_catalog.json # 全量教材元数据本地缓存(自动生成)
|
||||||
|
├── requirements.txt # 项目 Python 依赖清单
|
||||||
|
├── README.md # 项目使用说明文档
|
||||||
|
├── temp_pages/ # 图片下载临时缓存目录(下载后自动清理/保留)
|
||||||
|
└── downloads/ # 生成的高清 PDF 默认存放目录
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 🔍 技术原理解析
|
||||||
|
|
||||||
|
1. 略。我是不明白为什么免费教材免费提供阅读不提供免费下载。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## ⚠️ 免责声明 (Disclaimer)
|
||||||
|
|
||||||
|
1. 本项目仅供 Python 爬虫技术交流、逆向工程学习与个人学习研究使用,严禁用于任何商业用途或盈利活动。
|
||||||
|
2. 本项目下载的所有教材版权均归**人民教育出版社(PEP)**及相关版权所有方所有。
|
||||||
|
3. 使用本项目时请控制请求频率,严禁进行任何可能对官方服务器造成过大负载的行为。请于下载后 24 小时内自行删除,如需长期使用请购买或支持官方正版出版物。
|
||||||
|
4. 使用者因违反版权或不当使用造成的一切法律纠纷与责任,均由使用者个人自行承担,与本项目作者无关。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 📄 开源许可
|
||||||
|
|
||||||
|
本项目基于 [MIT License](LICENSE) 协议开源。欢迎提交 Issue 与 Pull Request!
|
||||||
|
|
||||||
@@ -54,12 +54,16 @@ def interactive_mode():
|
|||||||
print("\n请选择检索模式:")
|
print("\n请选择检索模式:")
|
||||||
print(" [1] 分类层级筛选(学段 ➔ 学科 ➔ 年级)")
|
print(" [1] 分类层级筛选(学段 ➔ 学科 ➔ 年级)")
|
||||||
print(" [2] 关键词全局搜索(如输入:'必修一'、'高一数学'、'生物')")
|
print(" [2] 关键词全局搜索(如输入:'必修一'、'高一数学'、'生物')")
|
||||||
|
print(" [3] 一键按「学段 ➔ 年级」两层目录全量下载全部教材")
|
||||||
|
|
||||||
mode_choice = input("请输入模式编号 (1/2, 默认 1): ").strip()
|
mode_choice = input("请输入模式编号 (1/2/3, 默认 1): ").strip()
|
||||||
|
|
||||||
matched_books = []
|
matched_books = []
|
||||||
|
|
||||||
if mode_choice == "2":
|
if mode_choice == "3":
|
||||||
|
matched_books = PepCatalog.fetch_and_decrypt_all()
|
||||||
|
print(f"\n[+] 已加载全网全部教材,共 {len(matched_books)} 本。")
|
||||||
|
elif mode_choice == "2":
|
||||||
kw = input("\n🔍 请输入搜索关键词: ").strip()
|
kw = input("\n🔍 请输入搜索关键词: ").strip()
|
||||||
if not kw:
|
if not kw:
|
||||||
print("[-] 关键词不能为空!")
|
print("[-] 关键词不能为空!")
|
||||||
@@ -93,13 +97,15 @@ def interactive_mode():
|
|||||||
|
|
||||||
print(f"\n✅ 共检索到 {len(matched_books)} 本教材:")
|
print(f"\n✅ 共检索到 {len(matched_books)} 本教材:")
|
||||||
print("-" * 65)
|
print("-" * 65)
|
||||||
for idx, b in enumerate(matched_books, 1):
|
for idx, b in enumerate(matched_books[:15], 1):
|
||||||
xd = b.get("xd", "")
|
xd = b.get("xd", "")
|
||||||
xk = b.get("xk", "")
|
xk = b.get("xk", "")
|
||||||
nj = b.get("nj", "")
|
nj = b.get("nj", "")
|
||||||
cc = b.get("cc", "")
|
cc = b.get("cc", "")
|
||||||
title = b.get("title", "")
|
title = b.get("title", "")
|
||||||
print(f" [{idx:2d}] [{xd}|{xk}|{nj}{cc}] 《{title}》 (ID: {b['id']})")
|
print(f" [{idx:2d}] [{xd}|{xk}|{nj}{cc}] 《{title}》 (ID: {b['id']})")
|
||||||
|
if len(matched_books) > 15:
|
||||||
|
print(f" ... 以及其余 {len(matched_books) - 15} 本教材(已折叠显示)")
|
||||||
print("-" * 65)
|
print("-" * 65)
|
||||||
|
|
||||||
print("\n请选择下载范围:")
|
print("\n请选择下载范围:")
|
||||||
@@ -139,14 +145,28 @@ def interactive_mode():
|
|||||||
return
|
return
|
||||||
|
|
||||||
out_dir = os.path.abspath("./downloads")
|
out_dir = os.path.abspath("./downloads")
|
||||||
print(f"\n🚀 即将开始下载 {len(to_download)} 本教材,保存目录: {out_dir}")
|
print(f"\n🚀 即将开始下载 {len(to_download)} 本教材,基础保存目录: {out_dir}")
|
||||||
|
print(f"🌲 采用「学段 ➔ 年级」两层子目录分类保存 (已存在 PDF 自动跳过)")
|
||||||
downloader = PepDownloader(headless=True, output_dir=out_dir)
|
downloader = PepDownloader(headless=True, output_dir=out_dir)
|
||||||
|
|
||||||
for idx, b in enumerate(to_download, 1):
|
for idx, b in enumerate(to_download, 1):
|
||||||
|
xd = b.get("xd", "其他学段")
|
||||||
|
nj = b.get("nj", "通用").strip() or "通用"
|
||||||
|
import re
|
||||||
|
safe_xd = re.sub(r'[\/:*?"<>|]', '_', xd).strip()
|
||||||
|
safe_nj = re.sub(r'[\/:*?"<>|]', '_', nj).strip()
|
||||||
|
sub_dir = os.path.join(safe_xd, safe_nj)
|
||||||
|
|
||||||
print(f"\n==================================================")
|
print(f"\n==================================================")
|
||||||
print(f"[{idx}/{len(to_download)}] 开始下载: 《{b.get('title')}》")
|
print(f"[{idx}/{len(to_download)}] [{safe_xd}/{safe_nj}] 《{b.get('title')}》")
|
||||||
print(f"==================================================")
|
print(f"==================================================")
|
||||||
downloader.download_book(book_id=b["id"], custom_title=b.get("title"))
|
downloader.download_book(
|
||||||
|
book_id=b["id"],
|
||||||
|
custom_title=b.get("title"),
|
||||||
|
sub_dir=sub_dir,
|
||||||
|
skip_if_exists=True,
|
||||||
|
clean_temp=True
|
||||||
|
)
|
||||||
|
|
||||||
print("\n🎉 全部选定任务执行完毕!")
|
print("\n🎉 全部选定任务执行完毕!")
|
||||||
|
|
||||||
@@ -154,14 +174,20 @@ def interactive_mode():
|
|||||||
def cli_args_mode(args):
|
def cli_args_mode(args):
|
||||||
"""命令行参数直接执行模式"""
|
"""命令行参数直接执行模式"""
|
||||||
print_banner()
|
print_banner()
|
||||||
|
if args.all:
|
||||||
|
matched = PepCatalog.fetch_and_decrypt_all()
|
||||||
|
else:
|
||||||
matched = PepCatalog.filter_books(xd=args.xd, xk=args.xk, nj=args.nj, keyword=args.search)
|
matched = PepCatalog.filter_books(xd=args.xd, xk=args.xk, nj=args.nj, keyword=args.search)
|
||||||
|
|
||||||
if not matched:
|
if not matched:
|
||||||
print("[-] 未查找到符合条件的教材!")
|
print("[-] 未查找到符合条件的教材!")
|
||||||
return
|
return
|
||||||
|
|
||||||
print(f"[+] 符合条件的教材共 {len(matched)} 本:")
|
print(f"[+] 符合条件的教材共 {len(matched)} 本:")
|
||||||
for idx, b in enumerate(matched, 1):
|
for idx, b in enumerate(matched[:15], 1):
|
||||||
print(f" [{idx}] 《{b.get('title')}》 (ID: {b['id']})")
|
print(f" [{idx}] [{b.get('xd')}|{b.get('nj')}] 《{b.get('title')}》 (ID: {b['id']})")
|
||||||
|
if len(matched) > 15:
|
||||||
|
print(f" ... 以及其余 {len(matched) - 15} 本教材(已折叠)")
|
||||||
|
|
||||||
if not args.yes:
|
if not args.yes:
|
||||||
confirm = input(f"\n确认下载以上 {len(matched)} 本教材吗?(y/n, 默认 y): ").strip().lower()
|
confirm = input(f"\n确认下载以上 {len(matched)} 本教材吗?(y/n, 默认 y): ").strip().lower()
|
||||||
@@ -171,20 +197,40 @@ def cli_args_mode(args):
|
|||||||
|
|
||||||
out_dir = os.path.abspath(args.output)
|
out_dir = os.path.abspath(args.output)
|
||||||
downloader = PepDownloader(headless=True, output_dir=out_dir)
|
downloader = PepDownloader(headless=True, output_dir=out_dir)
|
||||||
|
use_tree = not args.flat
|
||||||
|
|
||||||
|
print(f"\n📂 保存根目录: {out_dir}")
|
||||||
|
print(f"🌲 目录结构: {'按「学段/年级」两层子目录' if use_tree else '全部平铺在根目录'}")
|
||||||
|
|
||||||
for idx, b in enumerate(matched, 1):
|
for idx, b in enumerate(matched, 1):
|
||||||
|
sub_dir = None
|
||||||
|
if use_tree:
|
||||||
|
import re
|
||||||
|
safe_xd = re.sub(r'[\/:*?"<>|]', '_', b.get("xd", "其他学段")).strip()
|
||||||
|
safe_nj = re.sub(r'[\/:*?"<>|]', '_', b.get("nj", "通用") or "通用").strip()
|
||||||
|
sub_dir = os.path.join(safe_xd, safe_nj)
|
||||||
|
|
||||||
print(f"\n[{idx}/{len(matched)}] 正在下载: 《{b.get('title')}》...")
|
print(f"\n[{idx}/{len(matched)}] 正在下载: 《{b.get('title')}》...")
|
||||||
downloader.download_book(book_id=b["id"], custom_title=b.get("title"))
|
downloader.download_book(
|
||||||
|
book_id=b["id"],
|
||||||
|
custom_title=b.get("title"),
|
||||||
|
sub_dir=sub_dir,
|
||||||
|
skip_if_exists=True,
|
||||||
|
clean_temp=True
|
||||||
|
)
|
||||||
|
|
||||||
print(f"\n🎉 下载完成!文件已保存至: {out_dir}")
|
print(f"\n🎉 下载完成!文件已保存至: {out_dir}")
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
parser = argparse.ArgumentParser(description="人民教育出版社电子教材 CLI 下载器")
|
parser = argparse.ArgumentParser(description="人民教育出版社电子教材 CLI 下载器")
|
||||||
|
parser.add_argument("--all", "-a", action="store_true", help="下载全网全部 780+ 本教材")
|
||||||
parser.add_argument("--xd", help="指定学段(如:小学(六三学制)、初中(六三学制)、高中等)")
|
parser.add_argument("--xd", help="指定学段(如:小学(六三学制)、初中(六三学制)、高中等)")
|
||||||
parser.add_argument("--xk", help="指定学科(如:语文、数学、英语、物理等)")
|
parser.add_argument("--xk", help="指定学科(如:语文、数学、英语、物理等)")
|
||||||
parser.add_argument("--nj", help="指定年级(如:一年级、七年级、必修等)")
|
parser.add_argument("--nj", help="指定年级(如:一年级、七年级、必修等)")
|
||||||
parser.add_argument("--search", "-s", help="全局搜索关键词")
|
parser.add_argument("--search", "-s", help="全局搜索关键词")
|
||||||
parser.add_argument("--output", "-o", default="./downloads", help="PDF 文件保存目录 (默认: ./downloads)")
|
parser.add_argument("--output", "-o", default="./downloads", help="PDF 文件保存目录 (默认: ./downloads)")
|
||||||
|
parser.add_argument("--flat", action="store_true", help="平铺存放在根目录下(默认自动按「学段/年级」两层子目录分类)")
|
||||||
parser.add_argument("--yes", "-y", action="store_true", help="免确认直接开始下载")
|
parser.add_argument("--yes", "-y", action="store_true", help="免确认直接开始下载")
|
||||||
parser.add_argument("--refresh", "-r", action="store_true", help="强制从官方服务器重新拉取并解密最新教材目录")
|
parser.add_argument("--refresh", "-r", action="store_true", help="强制从官方服务器重新拉取并解密最新教材目录")
|
||||||
parser.add_argument("--clear-cache", "-c", action="store_true", help="清理本地临时下载缓存 (temp_pages)")
|
parser.add_argument("--clear-cache", "-c", action="store_true", help="清理本地临时下载缓存 (temp_pages)")
|
||||||
@@ -204,7 +250,7 @@ def main():
|
|||||||
print(f"[+] 目录同步成功!共获取到 {len(books)} 本教材。")
|
print(f"[+] 目录同步成功!共获取到 {len(books)} 本教材。")
|
||||||
return
|
return
|
||||||
|
|
||||||
if args.xd or args.xk or args.nj or args.search:
|
if args.all or args.xd or args.xk or args.nj or args.search:
|
||||||
cli_args_mode(args)
|
cli_args_mode(args)
|
||||||
else:
|
else:
|
||||||
interactive_mode()
|
interactive_mode()
|
||||||
|
|||||||
+233
@@ -0,0 +1,233 @@
|
|||||||
|
"""
|
||||||
|
人教社电子教材全量多线程下载器 (download_all.py)
|
||||||
|
功能:
|
||||||
|
按「学段 ➔ 年级」两层层级目录结构,多线程并发下载全网人教社电子教材并自动合成为高清 PDF。
|
||||||
|
|
||||||
|
默认目录结构示例:
|
||||||
|
downloads/
|
||||||
|
├── 小学(六三学制)/
|
||||||
|
│ ├── 一年级/
|
||||||
|
│ │ ├── 义务教育教科书 语文 一年级 上册.pdf
|
||||||
|
│ │ └── 义务教育教科书 数学 一年级 上册.pdf
|
||||||
|
│ └── 二年级/
|
||||||
|
├── 初中(六三学制)/
|
||||||
|
└── 高中/
|
||||||
|
└── 必修/
|
||||||
|
|
||||||
|
特性:
|
||||||
|
- 默认 3 线程并发加速下载(支持 --workers 自定义)
|
||||||
|
- 自动按「学段/年级」两层文件夹分类存放
|
||||||
|
- 紧凑输出:不显示单页下载细节,下载并合成完毕时直接提示
|
||||||
|
- 支持断点续传:已下载完成的教材自动秒跳过,无缝继续
|
||||||
|
- 自动清理单页切片图片缓存,极大节省磁盘空间
|
||||||
|
"""
|
||||||
|
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import re
|
||||||
|
import time
|
||||||
|
import random
|
||||||
|
import argparse
|
||||||
|
import threading
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
from pep_core import PepCatalog, PepDownloader, normalize_xd, get_base_dir
|
||||||
|
|
||||||
|
if hasattr(sys.stdout, "reconfigure"):
|
||||||
|
sys.stdout.reconfigure(encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def sanitize_filename(name: str) -> str:
|
||||||
|
"""过滤文件名中的非法字符"""
|
||||||
|
return re.sub(r'[\/:*?"<>|]', '_', name).strip()
|
||||||
|
|
||||||
|
|
||||||
|
def run_download_all(output_dir: str = "./downloads",
|
||||||
|
target_xd: str = None,
|
||||||
|
target_nj: str = None,
|
||||||
|
max_workers: int = 3,
|
||||||
|
delay: float = 1.0,
|
||||||
|
clean_temp: bool = True):
|
||||||
|
print("=" * 70)
|
||||||
|
print(" 📚 人民教育出版社 (PEP) 全量教材多线程层级下载器")
|
||||||
|
print("=" * 70)
|
||||||
|
|
||||||
|
# 1. 获取全量教材目录数据
|
||||||
|
print("[*] 正在加载教材全量目录数据...")
|
||||||
|
all_books = PepCatalog.fetch_and_decrypt_all()
|
||||||
|
print(f"[+] 成功获取全量教材数据库,共 {len(all_books)} 本。")
|
||||||
|
|
||||||
|
# 2. 条件过滤(支持多个学段,逗号分隔,如:"义务教育(六三学制),义务教育(五四学制),高中")
|
||||||
|
target_xds = []
|
||||||
|
if target_xd and target_xd != "全部":
|
||||||
|
target_xds = [x.strip() for x in re.split(r'[,,|/]', target_xd) if x.strip()]
|
||||||
|
|
||||||
|
def is_match_xd(b_xd, b_xdtype):
|
||||||
|
if not target_xds:
|
||||||
|
return True
|
||||||
|
for t in target_xds:
|
||||||
|
# 精确匹配或关键词包含匹配(如匹配 六三、五四、高中)
|
||||||
|
if t == b_xd or t == b_xdtype:
|
||||||
|
return True
|
||||||
|
if ("六三" in t and "六三" in b_xdtype) or ("六三" in t and "六三" in b_xd):
|
||||||
|
return True
|
||||||
|
if ("五四" in t or "五·四" in t) and ("五四" in b_xdtype or "五·四" in b_xdtype or "五四" in b_xd or "五·四" in b_xd):
|
||||||
|
return True
|
||||||
|
if ("高中" in t) and ("高中" in b_xd or "高中" in b_xdtype):
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
books_to_download = []
|
||||||
|
for b in all_books:
|
||||||
|
xd = normalize_xd(b.get("xd", "其他学段"))
|
||||||
|
xdtype = b.get("xdtype", "").strip()
|
||||||
|
nj = b.get("nj", "通用").strip() or "通用"
|
||||||
|
|
||||||
|
if not is_match_xd(xd, xdtype):
|
||||||
|
continue
|
||||||
|
if target_nj and target_nj != "全部" and nj != target_nj:
|
||||||
|
continue
|
||||||
|
books_to_download.append(b)
|
||||||
|
|
||||||
|
total_count = len(books_to_download)
|
||||||
|
if total_count == 0:
|
||||||
|
print("[-] 未查找到符合条件的教材!请检查学段或年级名称是否正确。")
|
||||||
|
return
|
||||||
|
|
||||||
|
base_out = os.path.abspath(output_dir) if output_dir else os.path.join(get_base_dir(), "downloads")
|
||||||
|
print(f"\n📂 基础保存目录: {base_out}")
|
||||||
|
print(f"📦 待处理教材数量: {total_count} 本")
|
||||||
|
print(f"🚀 并发下载线程数: {max_workers} 个线程")
|
||||||
|
print(f"🌲 归档目录结构: {base_out}/<学段>/<年级>/<教材名>.pdf")
|
||||||
|
print(f"⚡ 断点续传机制: 开启 (已存在 PDF 自动秒跳过)")
|
||||||
|
print(f"🧹 切片缓存清理: {'开启 (每本合成后自动删除临时图片)' if clean_temp else '关闭'}")
|
||||||
|
print("-" * 70 + "\n")
|
||||||
|
|
||||||
|
downloader = PepDownloader(headless=True, output_dir=base_out)
|
||||||
|
|
||||||
|
# 统计锁
|
||||||
|
lock = threading.Lock()
|
||||||
|
processed_count = 0
|
||||||
|
success_count = 0
|
||||||
|
skipped_count = 0
|
||||||
|
failed_count = 0
|
||||||
|
start_time = time.time()
|
||||||
|
|
||||||
|
def process_single_book(item):
|
||||||
|
nonlocal processed_count, success_count, skipped_count, failed_count
|
||||||
|
idx, b = item
|
||||||
|
book_id = b.get("id")
|
||||||
|
raw_title = b.get("title", f"book_{book_id}")
|
||||||
|
|
||||||
|
# 优先采用用户习惯的学段规范大类
|
||||||
|
xdtype = b.get("xdtype", "").strip()
|
||||||
|
xd = normalize_xd(b.get("xd", "其他学段"))
|
||||||
|
if "六三" in xdtype:
|
||||||
|
stage_dir_name = "义务教育(六三学制)"
|
||||||
|
elif "五四" in xdtype or "五·四" in xdtype:
|
||||||
|
stage_dir_name = "义务教育(五四学制)"
|
||||||
|
elif "高中" in xdtype or "高中" in xd:
|
||||||
|
stage_dir_name = "高中"
|
||||||
|
else:
|
||||||
|
stage_dir_name = xd
|
||||||
|
|
||||||
|
nj = b.get("nj", "通用").strip() or "通用"
|
||||||
|
|
||||||
|
safe_xd = sanitize_filename(stage_dir_name)
|
||||||
|
safe_nj = sanitize_filename(nj)
|
||||||
|
sub_dir = os.path.join(safe_xd, safe_nj)
|
||||||
|
safe_title = sanitize_filename(raw_title)
|
||||||
|
|
||||||
|
target_dir = os.path.join(base_out, sub_dir)
|
||||||
|
expected_pdf = os.path.join(target_dir, f"{safe_title}.pdf")
|
||||||
|
|
||||||
|
# 检查是否已存在完整 PDF
|
||||||
|
if os.path.exists(expected_pdf) and os.path.getsize(expected_pdf) > 50000:
|
||||||
|
with lock:
|
||||||
|
processed_count += 1
|
||||||
|
skipped_count += 1
|
||||||
|
percent = (processed_count / total_count) * 100
|
||||||
|
size_kb = os.path.getsize(expected_pdf) // 1024
|
||||||
|
print(f"[{processed_count}/{total_count}] ({percent:5.1f}%) [✔ 已存在跳过] [{safe_xd}/{safe_nj}] 《{safe_title}》 ({size_kb} KB)")
|
||||||
|
return True
|
||||||
|
|
||||||
|
# 开始下载
|
||||||
|
try:
|
||||||
|
# 错峰启动微延时
|
||||||
|
time.sleep(random.uniform(0.1, 0.8))
|
||||||
|
|
||||||
|
pdf_path = downloader.download_book(
|
||||||
|
book_id=book_id,
|
||||||
|
custom_title=raw_title,
|
||||||
|
sub_dir=sub_dir,
|
||||||
|
skip_if_exists=True,
|
||||||
|
clean_temp=clean_temp,
|
||||||
|
quiet=True
|
||||||
|
)
|
||||||
|
|
||||||
|
with lock:
|
||||||
|
processed_count += 1
|
||||||
|
percent = (processed_count / total_count) * 100
|
||||||
|
if pdf_path and os.path.exists(pdf_path):
|
||||||
|
success_count += 1
|
||||||
|
pdf_size_mb = round(os.path.getsize(pdf_path) / (1024 * 1024), 2)
|
||||||
|
print(f"[{processed_count}/{total_count}] ({percent:5.1f}%) [🎉 下载合成完成] [{safe_xd}/{safe_nj}] 《{safe_title}》 ({pdf_size_mb} MB)")
|
||||||
|
else:
|
||||||
|
failed_count += 1
|
||||||
|
print(f"[{processed_count}/{total_count}] ({percent:5.1f}%) [❌ 下载失败] [{safe_xd}/{safe_nj}] 《{safe_title}》")
|
||||||
|
|
||||||
|
if delay > 0:
|
||||||
|
time.sleep(random.uniform(delay * 0.7, delay * 1.3))
|
||||||
|
return True
|
||||||
|
except Exception as e:
|
||||||
|
with lock:
|
||||||
|
processed_count += 1
|
||||||
|
failed_count += 1
|
||||||
|
percent = (processed_count / total_count) * 100
|
||||||
|
print(f"[{processed_count}/{total_count}] ({percent:5.1f}%) [❌ 下载异常: {e}] [{safe_xd}/{safe_nj}] 《{safe_title}》")
|
||||||
|
return False
|
||||||
|
|
||||||
|
indexed_books = list(enumerate(books_to_download, 1))
|
||||||
|
|
||||||
|
# 使用线程池并发下载
|
||||||
|
with ThreadPoolExecutor(max_workers=max_workers) as executor:
|
||||||
|
futures = [executor.submit(process_single_book, item) for item in indexed_books]
|
||||||
|
for f in as_completed(futures):
|
||||||
|
try:
|
||||||
|
f.result()
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
|
||||||
|
elapsed = time.time() - start_time
|
||||||
|
elapsed_min = round(elapsed / 60, 1)
|
||||||
|
|
||||||
|
print("\n" + "=" * 70)
|
||||||
|
print("🎉 批量下载任务执行完毕!")
|
||||||
|
print(f"📊 统计汇总: 总计 {total_count} 本 | 新下载完成 {success_count} 本 | 已存在跳过 {skipped_count} 本 | 失败 {failed_count} 本")
|
||||||
|
print(f"⏱️ 总耗时: {elapsed_min} 分钟")
|
||||||
|
print(f"📁 教材归档存放于: {base_out}")
|
||||||
|
print("=" * 70)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
parser = argparse.ArgumentParser(description="人教社电子教材按「学段/年级」层级目录多线程全量下载工具")
|
||||||
|
parser.add_argument("--workers", "-w", type=int, default=3, help="同时下载的并发线程数 (默认: 3)")
|
||||||
|
parser.add_argument("--output", "-o", default="./downloads", help="PDF 根输出目录 (默认: ./downloads)")
|
||||||
|
parser.add_argument("--xd", help="只下载指定学段(如:小学(六三学制)、初中(六三学制)、高中 等)")
|
||||||
|
parser.add_argument("--nj", help="只下载指定年级(如:一年级、七年级 等)")
|
||||||
|
parser.add_argument("--delay", "-d", type=float, default=0.5, help="单线程任务间休眠秒数 (默认: 0.5 秒)")
|
||||||
|
parser.add_argument("--keep-temp", action="store_true", help="保留单页切片图片缓存(默认会自动删除以节约空间)")
|
||||||
|
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
run_download_all(
|
||||||
|
output_dir=args.output,
|
||||||
|
target_xd=args.xd,
|
||||||
|
target_nj=args.nj,
|
||||||
|
max_workers=args.workers,
|
||||||
|
delay=args.delay,
|
||||||
|
clean_temp=not args.keep_temp
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
+370
-496
File diff suppressed because it is too large
Load Diff
+128
-33
@@ -109,6 +109,36 @@ def sort_nj_key(nj: str) -> Tuple[int, int, str]:
|
|||||||
return (1, 0, nj)
|
return (1, 0, nj)
|
||||||
|
|
||||||
|
|
||||||
|
def map_book_xd(b: dict) -> str:
|
||||||
|
"""根据教材元数据中的 xd 与 xdtype 精确解析标准化学段"""
|
||||||
|
xd = b.get("xd", "").strip()
|
||||||
|
xdtype = b.get("xdtype", "").strip()
|
||||||
|
|
||||||
|
if "盲文" in xdtype:
|
||||||
|
return "盲校(盲文版)"
|
||||||
|
if "低视力" in xdtype:
|
||||||
|
return "盲校(低视力版)"
|
||||||
|
if "聋校" in xdtype:
|
||||||
|
return "聋校"
|
||||||
|
if "培智" in xdtype:
|
||||||
|
return "培智学校"
|
||||||
|
if "六三" in xdtype:
|
||||||
|
if xd == "初中":
|
||||||
|
return "初中(六三学制)"
|
||||||
|
return "小学(六三学制)"
|
||||||
|
if "五四" in xdtype:
|
||||||
|
if xd == "初中":
|
||||||
|
return "初中(五·四学制)"
|
||||||
|
return "小学(五四学制)"
|
||||||
|
if xd == "高中":
|
||||||
|
return "高中"
|
||||||
|
if xd == "小学":
|
||||||
|
return "小学(六三学制)"
|
||||||
|
if xd == "初中":
|
||||||
|
return "初中(六三学制)"
|
||||||
|
return normalize_xd(xd)
|
||||||
|
|
||||||
|
|
||||||
class PepCatalog:
|
class PepCatalog:
|
||||||
"""人教社教材目录管理器"""
|
"""人教社教材目录管理器"""
|
||||||
BASE_URL = "https://jc.pep.com.cn/"
|
BASE_URL = "https://jc.pep.com.cn/"
|
||||||
@@ -124,9 +154,8 @@ class PepCatalog:
|
|||||||
with open(cls.LOCAL_CACHE, "r", encoding="utf-8") as f:
|
with open(cls.LOCAL_CACHE, "r", encoding="utf-8") as f:
|
||||||
data = json.load(f)
|
data = json.load(f)
|
||||||
if isinstance(data, list) and len(data) > 0:
|
if isinstance(data, list) and len(data) > 0:
|
||||||
# 确保每条记录都包含标准化的学段
|
|
||||||
for b in data:
|
for b in data:
|
||||||
b["xd"] = normalize_xd(b.get("xd", ""))
|
b["xd"] = map_book_xd(b)
|
||||||
return data
|
return data
|
||||||
except Exception:
|
except Exception:
|
||||||
pass
|
pass
|
||||||
@@ -136,22 +165,21 @@ class PepCatalog:
|
|||||||
"Referer": cls.BASE_URL
|
"Referer": cls.BASE_URL
|
||||||
}
|
}
|
||||||
|
|
||||||
# 1. 下载首页定位 chunk-bfbdf2c4 JS
|
# 1. 下载首页定位 chunk-bfbdf2c4 JS (支持兼容压缩无引号属性)
|
||||||
req = urllib.request.Request(cls.BASE_URL, headers=headers)
|
req = urllib.request.Request(cls.BASE_URL, headers=headers)
|
||||||
with urllib.request.urlopen(req) as resp:
|
with urllib.request.urlopen(req) as resp:
|
||||||
html = resp.read().decode("utf-8", errors="ignore")
|
html = resp.read().decode("utf-8", errors="ignore")
|
||||||
|
|
||||||
chunk_match = re.search(r'src="(/js/chunk-bfbdf2c4\.[a-f0-9]+\.js)"', html)
|
chunk_match = re.search(r'(/js/chunk-bfbdf2c4\.[a-f0-9]+\.js)', html)
|
||||||
if not chunk_match:
|
chunk_path = chunk_match.group(1) if chunk_match else "/js/chunk-bfbdf2c4.b4dfb5d3.js"
|
||||||
chunk_match = re.search(r'href="(/js/chunk-bfbdf2c4\.[a-f0-9]+\.js)"', html)
|
|
||||||
|
|
||||||
chunk_path = chunk_match.group(1) if chunk_match else "/js/chunk-bfbdf2c4.3782cce3.js"
|
|
||||||
chunk_url = urllib.parse.urljoin(cls.BASE_URL, chunk_path)
|
chunk_url = urllib.parse.urljoin(cls.BASE_URL, chunk_path)
|
||||||
|
|
||||||
# 2. 获取 JS 内容并提取十六进制密文
|
# 2. 获取 JS 内容并提取十六进制密文
|
||||||
js_req = urllib.request.Request(chunk_url, headers=headers)
|
js_req = urllib.request.Request(chunk_url, headers=headers)
|
||||||
raw_js = urllib.request.urlopen(js_req).read()
|
raw_js = urllib.request.urlopen(js_req).read()
|
||||||
|
|
||||||
|
m = re.search(rb'c\s*=\s*"([A-F0-9]+)"', raw_js)
|
||||||
|
if not m:
|
||||||
m = re.search(rb'var\s+o,\s*c\s*=\s*"([A-F0-9]+)"', raw_js)
|
m = re.search(rb'var\s+o,\s*c\s*=\s*"([A-F0-9]+)"', raw_js)
|
||||||
if not m:
|
if not m:
|
||||||
raise ValueError("未能从前端 JS 中匹配到教材数据密文!")
|
raise ValueError("未能从前端 JS 中匹配到教材数据密文!")
|
||||||
@@ -170,9 +198,9 @@ class PepCatalog:
|
|||||||
data_json = json.loads(plaintext.decode("utf-8"))
|
data_json = json.loads(plaintext.decode("utf-8"))
|
||||||
items = data_json.get("data", [])
|
items = data_json.get("data", [])
|
||||||
|
|
||||||
# 标准化学段名称
|
# 精确标准化学段名称
|
||||||
for b in items:
|
for b in items:
|
||||||
b["xd"] = normalize_xd(b.get("xd", ""))
|
b["xd"] = map_book_xd(b)
|
||||||
|
|
||||||
# 缓存到本地
|
# 缓存到本地
|
||||||
with open(cls.LOCAL_CACHE, "w", encoding="utf-8") as f:
|
with open(cls.LOCAL_CACHE, "w", encoding="utf-8") as f:
|
||||||
@@ -284,17 +312,19 @@ class PepDownloader:
|
|||||||
os.makedirs(self.output_dir, exist_ok=True)
|
os.makedirs(self.output_dir, exist_ok=True)
|
||||||
|
|
||||||
def _solve_slider(self, page, log_cb: Optional[Callable[[str], None]] = None) -> bool:
|
def _solve_slider(self, page, log_cb: Optional[Callable[[str], None]] = None) -> bool:
|
||||||
"""检测并破解阿里云 WAF 滑块"""
|
"""检测并高可靠破解阿里云 WAF 滑块(支持自动重试与防伪装刷新)"""
|
||||||
|
for check_i in range(3):
|
||||||
try:
|
try:
|
||||||
page.wait_for_selector(".btn_slide", timeout=3000)
|
page.wait_for_selector(".btn_slide", timeout=2500)
|
||||||
except Exception:
|
except Exception:
|
||||||
pass
|
pass
|
||||||
|
|
||||||
slider = page.query_selector(".btn_slide")
|
slider = page.query_selector(".btn_slide")
|
||||||
if not slider:
|
if not slider:
|
||||||
|
# 若无滑块,直接返回通过
|
||||||
return True
|
return True
|
||||||
|
|
||||||
msg = "[*] 检测到阿里云 WAF 滑块验证码,正在自动滑动破解..."
|
msg = f"[*] 检测到阿里云 WAF 滑块验证码 (第 {check_i + 1} 次尝试),正在自动滑动破解..."
|
||||||
if log_cb: log_cb(msg)
|
if log_cb: log_cb(msg)
|
||||||
else: print(msg)
|
else: print(msg)
|
||||||
|
|
||||||
@@ -303,31 +333,50 @@ class PepDownloader:
|
|||||||
scale_box = scale.bounding_box() if scale else None
|
scale_box = scale.bounding_box() if scale else None
|
||||||
|
|
||||||
if not (box and scale_box):
|
if not (box and scale_box):
|
||||||
return False
|
time.sleep(1)
|
||||||
|
continue
|
||||||
|
|
||||||
start_x = box["x"] + box["width"] / 2
|
start_x = box["x"] + box["width"] / 2
|
||||||
start_y = box["y"] + box["height"] / 2
|
start_y = box["y"] + box["height"] / 2
|
||||||
distance = scale_box["width"] - box["width"] + 5
|
distance = scale_box["width"] - box["width"] + 5
|
||||||
target_x = start_x + distance
|
|
||||||
|
|
||||||
page.mouse.move(start_x, start_y)
|
page.mouse.move(start_x, start_y)
|
||||||
time.sleep(random.uniform(0.15, 0.3))
|
time.sleep(random.uniform(0.15, 0.25))
|
||||||
page.mouse.down()
|
page.mouse.down()
|
||||||
time.sleep(0.1)
|
time.sleep(0.05)
|
||||||
|
|
||||||
steps = random.randint(35, 45)
|
# 模拟高拟真人手轨迹:带初段加速与末端微晃动
|
||||||
|
steps = random.randint(28, 38)
|
||||||
for i in range(1, steps + 1):
|
for i in range(1, steps + 1):
|
||||||
t = i / steps
|
t = i / steps
|
||||||
progress = 1 - (1 - t) * (1 - t)
|
# 缓动函数
|
||||||
curr_x = start_x + distance * progress + random.uniform(-1, 1)
|
progress = math.sin(t * (math.pi / 2))
|
||||||
curr_y = start_y + math.sin(t * math.pi) * 2 + random.uniform(-1, 1)
|
curr_x = start_x + distance * progress + random.uniform(-0.8, 0.8)
|
||||||
|
curr_y = start_y + random.uniform(-1.2, 1.2)
|
||||||
page.mouse.move(curr_x, curr_y)
|
page.mouse.move(curr_x, curr_y)
|
||||||
time.sleep(random.uniform(0.01, 0.025))
|
time.sleep(random.uniform(0.012, 0.022))
|
||||||
|
|
||||||
time.sleep(0.1)
|
# 确保推到最右侧
|
||||||
|
page.mouse.move(start_x + distance + random.randint(2, 6), start_y)
|
||||||
|
time.sleep(0.08)
|
||||||
page.mouse.up()
|
page.mouse.up()
|
||||||
time.sleep(2.5)
|
time.sleep(2.5)
|
||||||
|
|
||||||
|
if page.query_selector(".btn_slide") is None:
|
||||||
|
res_msg = "[+] 滑块验证通过!"
|
||||||
|
if log_cb: log_cb(res_msg)
|
||||||
|
else: print(res_msg)
|
||||||
|
return True
|
||||||
|
else:
|
||||||
|
# 若滑块仍在,检查是否有“点击刷新”按钮
|
||||||
|
reload_btn = page.query_selector(".nc_iconfont.btn_refresh, .errloading a, .nc-lang-cnt a")
|
||||||
|
if reload_btn:
|
||||||
|
try:
|
||||||
|
reload_btn.click()
|
||||||
|
time.sleep(2)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
|
||||||
success = page.query_selector(".btn_slide") is None
|
success = page.query_selector(".btn_slide") is None
|
||||||
res_msg = "[+] 滑块验证通过!" if success else "[-] 滑块验证未通过。"
|
res_msg = "[+] 滑块验证通过!" if success else "[-] 滑块验证未通过。"
|
||||||
if log_cb: log_cb(res_msg)
|
if log_cb: log_cb(res_msg)
|
||||||
@@ -337,16 +386,38 @@ class PepDownloader:
|
|||||||
def download_book(self,
|
def download_book(self,
|
||||||
book_id: str,
|
book_id: str,
|
||||||
custom_title: Optional[str] = None,
|
custom_title: Optional[str] = None,
|
||||||
|
sub_dir: Optional[str] = None,
|
||||||
progress_cb: Optional[Callable[[int, int, str], None]] = None,
|
progress_cb: Optional[Callable[[int, int, str], None]] = None,
|
||||||
log_cb: Optional[Callable[[str], None]] = None) -> Optional[str]:
|
log_cb: Optional[Callable[[str], None]] = None,
|
||||||
|
skip_if_exists: bool = True,
|
||||||
|
clean_temp: bool = True,
|
||||||
|
quiet: bool = False) -> Optional[str]:
|
||||||
"""
|
"""
|
||||||
下载单本教材
|
下载单本教材
|
||||||
:param book_id: 教材 ID(如 1284001101241)
|
:param book_id: 教材 ID(如 1284001101241)
|
||||||
:param custom_title: 自定义书名
|
:param custom_title: 自定义书名
|
||||||
|
:param sub_dir: 子目录(如 "小学(六三学制)/一年级"),实现层级分类保存
|
||||||
:param progress_cb: 进度回调 (current_page, total_pages, status_text)
|
:param progress_cb: 进度回调 (current_page, total_pages, status_text)
|
||||||
:param log_cb: 日志回调 (log_text)
|
:param log_cb: 日志回调 (log_text)
|
||||||
|
:param skip_if_exists: 若本地已存在完整 PDF 则自动跳过
|
||||||
|
:param clean_temp: 合成 PDF 后自动删除该书的切片图片以节约磁盘空间
|
||||||
|
:param quiet: 静默模式,不打印各页下载过程,仅在关键节点或报错时提示
|
||||||
:return: 生成的 PDF 绝对路径
|
:return: 生成的 PDF 绝对路径
|
||||||
"""
|
"""
|
||||||
|
target_dir = os.path.join(self.output_dir, sub_dir) if sub_dir else self.output_dir
|
||||||
|
os.makedirs(target_dir, exist_ok=True)
|
||||||
|
|
||||||
|
# 检查是否已存在完整 PDF(断点续传/跳过机制)
|
||||||
|
if custom_title and skip_if_exists:
|
||||||
|
safe_title = re.sub(r'[\/:*?"<>|]', '_', custom_title).strip()
|
||||||
|
target_pdf = os.path.join(target_dir, f"{safe_title}.pdf")
|
||||||
|
if os.path.exists(target_pdf) and os.path.getsize(target_pdf) > 50000:
|
||||||
|
skip_msg = f"[✔] 本地已存在 《{safe_title}》 ({os.path.getsize(target_pdf) // 1024} KB),自动跳过。"
|
||||||
|
if log_cb: log_cb(skip_msg)
|
||||||
|
else: print(skip_msg)
|
||||||
|
if progress_cb: progress_cb(1, 1, "本地已存在,跳过")
|
||||||
|
return target_pdf
|
||||||
|
|
||||||
book_url = f"https://book.pep.com.cn/{book_id}/"
|
book_url = f"https://book.pep.com.cn/{book_id}/"
|
||||||
init_msg = f"[*] 准备加载教材: {book_url}"
|
init_msg = f"[*] 准备加载教材: {book_url}"
|
||||||
if log_cb: log_cb(init_msg)
|
if log_cb: log_cb(init_msg)
|
||||||
@@ -411,6 +482,15 @@ class PepDownloader:
|
|||||||
safe_title = re.sub(r'[\/:*?"<>|]', '_', final_title).strip()
|
safe_title = re.sub(r'[\/:*?"<>|]', '_', final_title).strip()
|
||||||
total_pages = config_info.get("total", 0)
|
total_pages = config_info.get("total", 0)
|
||||||
|
|
||||||
|
# 再次检查目标 PDF 是否存在
|
||||||
|
target_pdf = os.path.join(target_dir, f"{safe_title}.pdf")
|
||||||
|
if skip_if_exists and os.path.exists(target_pdf) and os.path.getsize(target_pdf) > 50000:
|
||||||
|
skip_msg = f"[✔] 本地已存在 《{safe_title}》 ({os.path.getsize(target_pdf) // 1024} KB),自动跳过。"
|
||||||
|
if log_cb: log_cb(skip_msg)
|
||||||
|
else: print(skip_msg)
|
||||||
|
browser.close()
|
||||||
|
return target_pdf
|
||||||
|
|
||||||
if total_pages == 0:
|
if total_pages == 0:
|
||||||
err = f"[-] 无法读取教材总页数 (URL: {page.url})"
|
err = f"[-] 无法读取教材总页数 (URL: {page.url})"
|
||||||
if log_cb: log_cb(err)
|
if log_cb: log_cb(err)
|
||||||
@@ -418,11 +498,12 @@ class PepDownloader:
|
|||||||
browser.close()
|
browser.close()
|
||||||
return None
|
return None
|
||||||
|
|
||||||
|
if not quiet:
|
||||||
info_msg = f"[+] 教材: 《{safe_title}》 | 总页数: {total_pages} 页"
|
info_msg = f"[+] 教材: 《{safe_title}》 | 总页数: {total_pages} 页"
|
||||||
if log_cb: log_cb(info_msg)
|
if log_cb: log_cb(info_msg)
|
||||||
else: print(info_msg)
|
else: print(info_msg)
|
||||||
|
|
||||||
temp_dir = os.path.join(".", "temp_pages", book_id)
|
temp_dir = os.path.join(get_base_dir(), "temp_pages", book_id)
|
||||||
os.makedirs(temp_dir, exist_ok=True)
|
os.makedirs(temp_dir, exist_ok=True)
|
||||||
image_files = []
|
image_files = []
|
||||||
|
|
||||||
@@ -439,7 +520,7 @@ class PepDownloader:
|
|||||||
continue
|
continue
|
||||||
|
|
||||||
download_success = False
|
download_success = False
|
||||||
for retry in range(3):
|
for retry in range(4):
|
||||||
res = page.evaluate("""async (url) => {
|
res = page.evaluate("""async (url) => {
|
||||||
try {
|
try {
|
||||||
const resp = await fetch(url);
|
const resp = await fetch(url);
|
||||||
@@ -467,6 +548,7 @@ class PepDownloader:
|
|||||||
f.write(raw_data)
|
f.write(raw_data)
|
||||||
image_files.append(img_path)
|
image_files.append(img_path)
|
||||||
size_kb = len(raw_data) // 1024
|
size_kb = len(raw_data) // 1024
|
||||||
|
if not quiet:
|
||||||
log_text = f" -> [{page_num}/{total_pages}] 下载成功 ({size_kb} KB)"
|
log_text = f" -> [{page_num}/{total_pages}] 下载成功 ({size_kb} KB)"
|
||||||
if log_cb: log_cb(log_text)
|
if log_cb: log_cb(log_text)
|
||||||
else: print(log_text)
|
else: print(log_text)
|
||||||
@@ -474,28 +556,34 @@ class PepDownloader:
|
|||||||
download_success = True
|
download_success = True
|
||||||
break
|
break
|
||||||
|
|
||||||
# 触发 WAF 验证码拦截
|
# 触发 WAF 验证码或网络抖动拦截
|
||||||
warn_msg = f" [!] 第 {page_num} 页触发验证码,自动刷新破解..."
|
warn_msg = f" [!] 《{safe_title}》第 {page_num} 页触发验证码/拦截 (重试 {retry + 1}/4),自动重新过盾..."
|
||||||
if log_cb: log_cb(warn_msg)
|
if log_cb: log_cb(warn_msg)
|
||||||
else: print(warn_msg)
|
else: print(warn_msg)
|
||||||
page.goto(book_url, referer="https://jc.pep.com.cn/", wait_until="networkidle")
|
|
||||||
time.sleep(1.5)
|
# 重新刷新/导航并破解滑块
|
||||||
|
try:
|
||||||
|
page.goto(book_url, referer="https://jc.pep.com.cn/", wait_until="load")
|
||||||
|
time.sleep(2)
|
||||||
self._solve_slider(page, log_cb)
|
self._solve_slider(page, log_cb)
|
||||||
|
time.sleep(2.5)
|
||||||
|
except Exception:
|
||||||
time.sleep(3)
|
time.sleep(3)
|
||||||
|
|
||||||
if not download_success:
|
if not download_success:
|
||||||
fail_msg = f" -> [{page_num}/{total_pages}] 下载失败!"
|
fail_msg = f" -> 《{safe_title}》[{page_num}/{total_pages}] 下载失败!"
|
||||||
if log_cb: log_cb(fail_msg)
|
if log_cb: log_cb(fail_msg)
|
||||||
else: print(fail_msg)
|
else: print(fail_msg)
|
||||||
|
|
||||||
time.sleep(random.uniform(0.35, 0.65))
|
time.sleep(random.uniform(0.15, 0.35))
|
||||||
|
|
||||||
browser.close()
|
browser.close()
|
||||||
|
|
||||||
if not image_files:
|
if not image_files:
|
||||||
return None
|
return None
|
||||||
|
|
||||||
output_pdf = os.path.join(self.output_dir, f"{safe_title}.pdf")
|
output_pdf = target_pdf
|
||||||
|
if not quiet:
|
||||||
merge_msg = f"[*] 正在合成 PDF: {os.path.basename(output_pdf)} ..."
|
merge_msg = f"[*] 正在合成 PDF: {os.path.basename(output_pdf)} ..."
|
||||||
if log_cb: log_cb(merge_msg)
|
if log_cb: log_cb(merge_msg)
|
||||||
else: print(merge_msg)
|
else: print(merge_msg)
|
||||||
@@ -518,6 +606,13 @@ class PepDownloader:
|
|||||||
rest_images = pil_images[1:]
|
rest_images = pil_images[1:]
|
||||||
first_im.save(output_pdf, "PDF", resolution=100.0, save_all=True, append_images=rest_images)
|
first_im.save(output_pdf, "PDF", resolution=100.0, save_all=True, append_images=rest_images)
|
||||||
|
|
||||||
|
# 下载合成完毕后,自动清理单页图片切片,释放磁盘空间
|
||||||
|
if clean_temp and os.path.exists(temp_dir):
|
||||||
|
try:
|
||||||
|
shutil.rmtree(temp_dir)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
|
||||||
done_msg = f"[✔] PDF 生成成功: {output_pdf}"
|
done_msg = f"[✔] PDF 生成成功: {output_pdf}"
|
||||||
if log_cb: log_cb(done_msg)
|
if log_cb: log_cb(done_msg)
|
||||||
else: print(done_msg)
|
else: print(done_msg)
|
||||||
|
|||||||
@@ -17,7 +17,7 @@ from fastapi import FastAPI, BackgroundTasks, WebSocket, WebSocketDisconnect
|
|||||||
from fastapi.responses import HTMLResponse, JSONResponse
|
from fastapi.responses import HTMLResponse, JSONResponse
|
||||||
from fastapi.staticfiles import StaticFiles
|
from fastapi.staticfiles import StaticFiles
|
||||||
|
|
||||||
from pep_core import PepCatalog, PepDownloader, XD_ORDER, XK_ORDER_PREFIX, NJ_ORDER, get_base_dir
|
from pep_core import PepCatalog, PepDownloader, XD_ORDER, XK_ORDER_PREFIX, NJ_ORDER, get_base_dir, normalize_xd
|
||||||
|
|
||||||
if hasattr(sys.stdout, "reconfigure"):
|
if hasattr(sys.stdout, "reconfigure"):
|
||||||
sys.stdout.reconfigure(encoding="utf-8")
|
sys.stdout.reconfigure(encoding="utf-8")
|
||||||
@@ -104,13 +104,22 @@ class TaskManager:
|
|||||||
self.log(txt)
|
self.log(txt)
|
||||||
|
|
||||||
try:
|
try:
|
||||||
|
xd = normalize_xd(book_to_download.get("xd", "其他学段"))
|
||||||
|
nj = book_to_download.get("nj", "通用").strip() or "通用"
|
||||||
|
safe_xd = re.sub(r'[\/:*?"<>|]', '_', xd).strip()
|
||||||
|
safe_nj = re.sub(r'[\/:*?"<>|]', '_', nj).strip()
|
||||||
|
sub_dir = os.path.join(safe_xd, safe_nj)
|
||||||
|
|
||||||
self.log(f"==================================================")
|
self.log(f"==================================================")
|
||||||
self.log(f"[*] 开始下载教材: 《{book_to_download.get('title')}》")
|
self.log(f"[*] 开始下载教材: [{safe_xd}/{safe_nj}] 《{book_to_download.get('title')}》")
|
||||||
downloader.download_book(
|
downloader.download_book(
|
||||||
book_id=book_to_download["id"],
|
book_id=book_to_download["id"],
|
||||||
custom_title=book_to_download.get("title"),
|
custom_title=book_to_download.get("title"),
|
||||||
|
sub_dir=sub_dir,
|
||||||
progress_cb=progress_callback,
|
progress_cb=progress_callback,
|
||||||
log_cb=log_callback
|
log_cb=log_callback,
|
||||||
|
skip_if_exists=True,
|
||||||
|
clean_temp=True
|
||||||
)
|
)
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
self.log(f"[-] 下载异常: {e}")
|
self.log(f"[-] 下载异常: {e}")
|
||||||
|
|||||||
Reference in New Issue
Block a user