v1.1 加入了命令行下的批量下载功能

现在可以一次把人教社的身体掏空了,实测最大10线程没问题。
This commit is contained in:
siknet
2026-09-04 15:43:36 +08:00
committed by GitHub
parent a3b03e3a64
commit 0b063467fd
6 changed files with 1043 additions and 577 deletions
@@ -0,0 +1,209 @@
# FreePEP 📚 人教社中小学电子教材批量下载器
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
**FreePEP** 是一款专为[人民教育出版社中小学电子教材平台](https://jc.pep.com.cn/)开发的自动化教材解析、批量抓取与高清 PDF 合成工具。
提供**WebUI 界面**与**交互式命令行**,内置全量 780+ 本教材目录(数据截止到2026年08月31日)自动解密引擎与阿里云 WAF 滑块验证码自动破解机制,支持一键下载指定学段、学科、年级的全套教材并自动生成高清 PDF 文件。
**更新内容:**
2026.09.04 FreePEP v1.1 命令行加入了批量下载功能,详情见**启动办法三**
---
## ✨ 核心特性
- 🎯 **双操作模式**:
- **现代化 WebUI 界面**:全响应式布局,还原官网层级筛选体验,支持复选框一键批量下载、实时进度条展示与本地保存目录快捷打开。
- **交互式 CLI 终端**:支持数字菜单导航选择与命令行参数直达(脚本集成与服务器环境友好)。
- 📦 **高清单页下载与 PDF 合成**:自动探测各教材实际总页数,批量下载高分辨率原始 JPG,并通过 Pillow 库自动合成为标准 PDF 文件。
---
## 📸 界面预览
### WebUI 网页端
![](https://img.wemd.app/1788139668274_4n5d28.png)
---
### CLI 网页端
![](https://img.wemd.app/1788139734008_2tmuak.png)
# 🚀 使用方式
## 🖥️ 方式一:(推荐,适合小白)
直接下载发行包,解压缩后运行FreePEP.exe,在弹出的Web页面里面自行操作下载。
## 方式二:从源码启动
### 克隆仓库与安装依赖
```bash
# 克隆仓库
git clone https://github.com/siknet/FreePEP.git
cd FreePEP
# 安装 Python 依赖
pip install -r requirements.txt
# 安装 Playwright 所需的 Chromium 浏览器内核
playwright install chromium
```
### 启动办法一:Webui
在终端中执行以下命令,系统将自动启动本地服务并在默认浏览器中打开管理页面:
```bash
python webui.py
```
- **访问地址**:`http://127.0.0.1:8000`
- **默认下载目录**:项目根目录下的 `./downloads` 文件夹(可在 Web 界面点击「📁 打开保存目录」直达)。
---
### 启动办法二:使用 CLI 命令行
#### 1. 交互式菜单模式
直接运行 `cli.py`,根据控制台提示逐步选择学段、学科与年级:
```bash
python cli.py
```
#### 2. 参数直达模式(适合自动化与脚本调用)
通过命令行参数直接指定条件过滤并自动开始下载:
```bash
# 1. 下载小学一年级的所有学科教材
python cli.py --xd "小学" --nj "一年级" -y
# 2. 下载高中数学的所有必修/选修教材
python cli.py --xd "高中" --xk "数学" -y
# 3. 下载初中全部教材
python cli.py --xd "初中" -y
# 4. 全局关键词搜索并下载(如包含"物理"的所有教材)
python cli.py --search "物理"
# 5. 清理本地临时图片缓存
python cli.py --clear-cache
# 6. 自定义 PDF 输出路径
python cli.py --xd "小学(六三学制)" --xk "语文" --nj "一年级" -o "D:/Textbooks" -y
```
**参数说明**:
| 参数 | 说明 | 示例 |
| :--- | :--- | :--- |
| `--xd` | 指定学段 | `小学(六三学制)`、`初中(六三学制)`、`小学(五四学制)`、`初中(五·四学制)`、`高中`、`培智学校`、`聋校`、`盲校(盲文版)`、`盲校(低视力版)` |
| `--xk` | 指定学科 | `语文`、`数学`、`英语`、`物理`、`化学`、`历史`、`道德与法治` 等 |
| `--nj` | 指定年级/册次 | `一年级`、`二年级` ... `九年级`、`专项`、`必修` 等 |
| `--search`, `-s` | 关键词全局搜索 | `必修`、`高一`、`地理` |
| `--clear-cache`, `-c` | 清理本地下载临时图片缓存 (`temp_pages`) | 无需参数 |
| `--refresh`, `-r` | 强制重新从官方服务器拉取解密最新目录 | 无需参数 |
| `--output`, `-o` | 指定 PDF 保存目录 | 默认: `./downloads` |
| `--yes`, `-y` | 跳过确认提示直接开始下载 | 开启免交互 |
### 启动办法三:按「学段 ➔ 年级」层级批量多线程下载 (download_all.py)
如果您希望将教材按照 **`学段/年级/`** 两层规范文件夹分类归档下载,直接运行专属的高速批量多线程下载脚本:
```bash
# 1. 一键下载全网全量教材(默认 3 线程并发,按「学段/年级」两层子目录自动归档)
python download_all.py
# 2. 启用多线程并发加速(例如 10 线程并行高速下载)
python download_all.py -w 10
# 3. 下载多个指定学段(支持逗号分隔,如:义务教育六三学制、五四学制与高中)并自定义保存路径
python download_all.py --xd "义务教育(六三学制),义务教育(五四学制),高中" -w 10 -o "D:/人教社教材"
# 4. 仅下载单个学段(如高中全部教材)
python download_all.py --xd "高中"
# 5. 仅下载某个指定年级教材
python download_all.py --nj "一年级"
```
**功能特性**:
- 🚀 **多线程并发提速**:默认 **3 线程**并发同时下载,支持通过 `-w / --workers` 自定义并发线程数(如 5~10 线程),下载效率大幅提升。
- 🎯 **支持多学段组合**:`--xd` 支持传入多个学段(逗号分隔),精准下载普通义务教育或高中学段,自动跳过特教教材。
- 🌲 **两层层级目录**:自动分类保存为 `downloads/<学段>/<年级>/<教材名>.pdf`,告别成百上千文件堆在单目录。
- ⚡ **智能断点续传**:已完整下载的教材**自动秒跳过**,中途随时中断无缝继续,绝不重复下载。
- 🧹 **切片自动清理**:每本教材合成 PDF 后自动清理临时图片分片,极大保护硬盘空间。
- 💬 **清爽紧凑输出**:不打印繁琐单页过程,仅在教材下载合并完成或跳过时输出进度与提示。
---
## 📦 一键打包为 Windows 独立 EXE 发行版
如果您想将本项目打包成 **脱离 Python 环境** 的独立绿色软件包发布给普通用户,只需执行工作区自带的一键打包脚本:
```bash
python build_exe.py
```
### 打包脚本会自动完成以下操作:
1. 自动检查并安装 `PyInstaller` 编译工具。
2. 将 WebUI 及所有 Python 依赖打包为独立可执行文件 `FreePEP.exe`。
3. **自动提取并内嵌绿色便携版 Chromium 浏览器内核**至 `browsers/` 目录。
4. 自动在 `dist/` 目录下生成 `FreePEP-Windows-x64.zip` 发行压缩包。
> **分发给用户使用**:用户下载压缩包后解压,**双击 `FreePEP.exe` 即可直接使用**(会自动弹出系统默认浏览器打开 WebUI,无需安装 Python、无需配置环境变量、无需额外下载浏览器内核)。
---
## 📁 项目目录结构
```text
FreePEP/
├── pep_core.py # 核心底层库(AES 解密、Playwright 爬取、PDF 合成)
├── webui.py # FastAPI WebUI 服务器与一体化前端界面
├── cli.py # 交互式与参数化 CLI 终端下载器
├── download_all.py # 按「学段/年级」层级全量下载脚本(支持断点续传与缓存清理)
├── build_exe.py # 一键打包发布 Windows EXE 独立便携包脚本
├── pep_crawler.py # 命令行测试与示例下载脚本
├── pep_catalog.json # 全量教材元数据本地缓存(自动生成)
├── requirements.txt # 项目 Python 依赖清单
├── README.md # 项目使用说明文档
├── temp_pages/ # 图片下载临时缓存目录(下载后自动清理/保留)
└── downloads/ # 生成的高清 PDF 默认存放目录
```
---
## 🔍 技术原理解析
1. 略。我是不明白为什么免费教材免费提供阅读不提供免费下载。
---
## ⚠️ 免责声明 (Disclaimer)
1. 本项目仅供 Python 爬虫技术交流、逆向工程学习与个人学习研究使用,严禁用于任何商业用途或盈利活动。
2. 本项目下载的所有教材版权均归**人民教育出版社(PEP)**及相关版权所有方所有。
3. 使用本项目时请控制请求频率,严禁进行任何可能对官方服务器造成过大负载的行为。请于下载后 24 小时内自行删除,如需长期使用请购买或支持官方正版出版物。
4. 使用者因违反版权或不当使用造成的一切法律纠纷与责任,均由使用者个人自行承担,与本项目作者无关。
---
## 📄 开源许可
本项目基于 [MIT License](LICENSE) 协议开源。欢迎提交 Issue 与 Pull Request!
+56 -10
View File
@@ -54,12 +54,16 @@ def interactive_mode():
print("\n请选择检索模式:") print("\n请选择检索模式:")
print(" [1] 分类层级筛选(学段 ➔ 学科 ➔ 年级)") print(" [1] 分类层级筛选(学段 ➔ 学科 ➔ 年级)")
print(" [2] 关键词全局搜索(如输入:'必修一'、'高一数学'、'生物')") print(" [2] 关键词全局搜索(如输入:'必修一'、'高一数学'、'生物')")
print(" [3] 一键按「学段 ➔ 年级」两层目录全量下载全部教材")
mode_choice = input("请输入模式编号 (1/2, 默认 1): ").strip() mode_choice = input("请输入模式编号 (1/2/3, 默认 1): ").strip()
matched_books = [] matched_books = []
if mode_choice == "2": if mode_choice == "3":
matched_books = PepCatalog.fetch_and_decrypt_all()
print(f"\n[+] 已加载全网全部教材,共 {len(matched_books)} 本。")
elif mode_choice == "2":
kw = input("\n🔍 请输入搜索关键词: ").strip() kw = input("\n🔍 请输入搜索关键词: ").strip()
if not kw: if not kw:
print("[-] 关键词不能为空!") print("[-] 关键词不能为空!")
@@ -93,13 +97,15 @@ def interactive_mode():
print(f"\n✅ 共检索到 {len(matched_books)} 本教材:") print(f"\n✅ 共检索到 {len(matched_books)} 本教材:")
print("-" * 65) print("-" * 65)
for idx, b in enumerate(matched_books, 1): for idx, b in enumerate(matched_books[:15], 1):
xd = b.get("xd", "") xd = b.get("xd", "")
xk = b.get("xk", "") xk = b.get("xk", "")
nj = b.get("nj", "") nj = b.get("nj", "")
cc = b.get("cc", "") cc = b.get("cc", "")
title = b.get("title", "") title = b.get("title", "")
print(f" [{idx:2d}] [{xd}|{xk}|{nj}{cc}] 《{title}》 (ID: {b['id']})") print(f" [{idx:2d}] [{xd}|{xk}|{nj}{cc}] 《{title}》 (ID: {b['id']})")
if len(matched_books) > 15:
print(f" ... 以及其余 {len(matched_books) - 15} 本教材(已折叠显示)")
print("-" * 65) print("-" * 65)
print("\n请选择下载范围:") print("\n请选择下载范围:")
@@ -139,14 +145,28 @@ def interactive_mode():
return return
out_dir = os.path.abspath("./downloads") out_dir = os.path.abspath("./downloads")
print(f"\n🚀 即将开始下载 {len(to_download)} 本教材,保存目录: {out_dir}") print(f"\n🚀 即将开始下载 {len(to_download)} 本教材,基础保存目录: {out_dir}")
print(f"🌲 采用「学段 ➔ 年级」两层子目录分类保存 (已存在 PDF 自动跳过)")
downloader = PepDownloader(headless=True, output_dir=out_dir) downloader = PepDownloader(headless=True, output_dir=out_dir)
for idx, b in enumerate(to_download, 1): for idx, b in enumerate(to_download, 1):
xd = b.get("xd", "其他学段")
nj = b.get("nj", "通用").strip() or "通用"
import re
safe_xd = re.sub(r'[\/:*?"<>|]', '_', xd).strip()
safe_nj = re.sub(r'[\/:*?"<>|]', '_', nj).strip()
sub_dir = os.path.join(safe_xd, safe_nj)
print(f"\n==================================================") print(f"\n==================================================")
print(f"[{idx}/{len(to_download)}] 开始下载: 《{b.get('title')}》") print(f"[{idx}/{len(to_download)}] [{safe_xd}/{safe_nj}] 《{b.get('title')}》")
print(f"==================================================") print(f"==================================================")
downloader.download_book(book_id=b["id"], custom_title=b.get("title")) downloader.download_book(
book_id=b["id"],
custom_title=b.get("title"),
sub_dir=sub_dir,
skip_if_exists=True,
clean_temp=True
)
print("\n🎉 全部选定任务执行完毕!") print("\n🎉 全部选定任务执行完毕!")
@@ -154,14 +174,20 @@ def interactive_mode():
def cli_args_mode(args): def cli_args_mode(args):
"""命令行参数直接执行模式""" """命令行参数直接执行模式"""
print_banner() print_banner()
if args.all:
matched = PepCatalog.fetch_and_decrypt_all()
else:
matched = PepCatalog.filter_books(xd=args.xd, xk=args.xk, nj=args.nj, keyword=args.search) matched = PepCatalog.filter_books(xd=args.xd, xk=args.xk, nj=args.nj, keyword=args.search)
if not matched: if not matched:
print("[-] 未查找到符合条件的教材!") print("[-] 未查找到符合条件的教材!")
return return
print(f"[+] 符合条件的教材共 {len(matched)} 本:") print(f"[+] 符合条件的教材共 {len(matched)} 本:")
for idx, b in enumerate(matched, 1): for idx, b in enumerate(matched[:15], 1):
print(f" [{idx}] 《{b.get('title')}》 (ID: {b['id']})") print(f" [{idx}] [{b.get('xd')}|{b.get('nj')}] 《{b.get('title')}》 (ID: {b['id']})")
if len(matched) > 15:
print(f" ... 以及其余 {len(matched) - 15} 本教材(已折叠)")
if not args.yes: if not args.yes:
confirm = input(f"\n确认下载以上 {len(matched)} 本教材吗?(y/n, 默认 y): ").strip().lower() confirm = input(f"\n确认下载以上 {len(matched)} 本教材吗?(y/n, 默认 y): ").strip().lower()
@@ -171,20 +197,40 @@ def cli_args_mode(args):
out_dir = os.path.abspath(args.output) out_dir = os.path.abspath(args.output)
downloader = PepDownloader(headless=True, output_dir=out_dir) downloader = PepDownloader(headless=True, output_dir=out_dir)
use_tree = not args.flat
print(f"\n📂 保存根目录: {out_dir}")
print(f"🌲 目录结构: {'按「学段/年级」两层子目录' if use_tree else '全部平铺在根目录'}")
for idx, b in enumerate(matched, 1): for idx, b in enumerate(matched, 1):
sub_dir = None
if use_tree:
import re
safe_xd = re.sub(r'[\/:*?"<>|]', '_', b.get("xd", "其他学段")).strip()
safe_nj = re.sub(r'[\/:*?"<>|]', '_', b.get("nj", "通用") or "通用").strip()
sub_dir = os.path.join(safe_xd, safe_nj)
print(f"\n[{idx}/{len(matched)}] 正在下载: 《{b.get('title')}》...") print(f"\n[{idx}/{len(matched)}] 正在下载: 《{b.get('title')}》...")
downloader.download_book(book_id=b["id"], custom_title=b.get("title")) downloader.download_book(
book_id=b["id"],
custom_title=b.get("title"),
sub_dir=sub_dir,
skip_if_exists=True,
clean_temp=True
)
print(f"\n🎉 下载完成!文件已保存至: {out_dir}") print(f"\n🎉 下载完成!文件已保存至: {out_dir}")
def main(): def main():
parser = argparse.ArgumentParser(description="人民教育出版社电子教材 CLI 下载器") parser = argparse.ArgumentParser(description="人民教育出版社电子教材 CLI 下载器")
parser.add_argument("--all", "-a", action="store_true", help="下载全网全部 780+ 本教材")
parser.add_argument("--xd", help="指定学段(如:小学(六三学制)、初中(六三学制)、高中等)") parser.add_argument("--xd", help="指定学段(如:小学(六三学制)、初中(六三学制)、高中等)")
parser.add_argument("--xk", help="指定学科(如:语文、数学、英语、物理等)") parser.add_argument("--xk", help="指定学科(如:语文、数学、英语、物理等)")
parser.add_argument("--nj", help="指定年级(如:一年级、七年级、必修等)") parser.add_argument("--nj", help="指定年级(如:一年级、七年级、必修等)")
parser.add_argument("--search", "-s", help="全局搜索关键词") parser.add_argument("--search", "-s", help="全局搜索关键词")
parser.add_argument("--output", "-o", default="./downloads", help="PDF 文件保存目录 (默认: ./downloads)") parser.add_argument("--output", "-o", default="./downloads", help="PDF 文件保存目录 (默认: ./downloads)")
parser.add_argument("--flat", action="store_true", help="平铺存放在根目录下(默认自动按「学段/年级」两层子目录分类)")
parser.add_argument("--yes", "-y", action="store_true", help="免确认直接开始下载") parser.add_argument("--yes", "-y", action="store_true", help="免确认直接开始下载")
parser.add_argument("--refresh", "-r", action="store_true", help="强制从官方服务器重新拉取并解密最新教材目录") parser.add_argument("--refresh", "-r", action="store_true", help="强制从官方服务器重新拉取并解密最新教材目录")
parser.add_argument("--clear-cache", "-c", action="store_true", help="清理本地临时下载缓存 (temp_pages)") parser.add_argument("--clear-cache", "-c", action="store_true", help="清理本地临时下载缓存 (temp_pages)")
@@ -204,7 +250,7 @@ def main():
print(f"[+] 目录同步成功!共获取到 {len(books)} 本教材。") print(f"[+] 目录同步成功!共获取到 {len(books)} 本教材。")
return return
if args.xd or args.xk or args.nj or args.search: if args.all or args.xd or args.xk or args.nj or args.search:
cli_args_mode(args) cli_args_mode(args)
else: else:
interactive_mode() interactive_mode()
+233
View File
@@ -0,0 +1,233 @@
"""
人教社电子教材全量多线程下载器 (download_all.py)
功能:
按「学段 ➔ 年级」两层层级目录结构,多线程并发下载全网人教社电子教材并自动合成为高清 PDF。
默认目录结构示例:
downloads/
├── 小学(六三学制)/
│ ├── 一年级/
│ │ ├── 义务教育教科书 语文 一年级 上册.pdf
│ │ └── 义务教育教科书 数学 一年级 上册.pdf
│ └── 二年级/
├── 初中(六三学制)/
└── 高中/
└── 必修/
特性:
- 默认 3 线程并发加速下载(支持 --workers 自定义)
- 自动按「学段/年级」两层文件夹分类存放
- 紧凑输出:不显示单页下载细节,下载并合成完毕时直接提示
- 支持断点续传:已下载完成的教材自动秒跳过,无缝继续
- 自动清理单页切片图片缓存,极大节省磁盘空间
"""
import os
import sys
import re
import time
import random
import argparse
import threading
from concurrent.futures import ThreadPoolExecutor, as_completed
from pep_core import PepCatalog, PepDownloader, normalize_xd, get_base_dir
if hasattr(sys.stdout, "reconfigure"):
sys.stdout.reconfigure(encoding="utf-8")
def sanitize_filename(name: str) -> str:
"""过滤文件名中的非法字符"""
return re.sub(r'[\/:*?"<>|]', '_', name).strip()
def run_download_all(output_dir: str = "./downloads",
target_xd: str = None,
target_nj: str = None,
max_workers: int = 3,
delay: float = 1.0,
clean_temp: bool = True):
print("=" * 70)
print(" 📚 人民教育出版社 (PEP) 全量教材多线程层级下载器")
print("=" * 70)
# 1. 获取全量教材目录数据
print("[*] 正在加载教材全量目录数据...")
all_books = PepCatalog.fetch_and_decrypt_all()
print(f"[+] 成功获取全量教材数据库,共 {len(all_books)} 本。")
# 2. 条件过滤(支持多个学段,逗号分隔,如:"义务教育(六三学制),义务教育(五四学制),高中")
target_xds = []
if target_xd and target_xd != "全部":
target_xds = [x.strip() for x in re.split(r'[,,|/]', target_xd) if x.strip()]
def is_match_xd(b_xd, b_xdtype):
if not target_xds:
return True
for t in target_xds:
# 精确匹配或关键词包含匹配(如匹配 六三、五四、高中)
if t == b_xd or t == b_xdtype:
return True
if ("六三" in t and "六三" in b_xdtype) or ("六三" in t and "六三" in b_xd):
return True
if ("五四" in t or "五·四" in t) and ("五四" in b_xdtype or "五·四" in b_xdtype or "五四" in b_xd or "五·四" in b_xd):
return True
if ("高中" in t) and ("高中" in b_xd or "高中" in b_xdtype):
return True
return False
books_to_download = []
for b in all_books:
xd = normalize_xd(b.get("xd", "其他学段"))
xdtype = b.get("xdtype", "").strip()
nj = b.get("nj", "通用").strip() or "通用"
if not is_match_xd(xd, xdtype):
continue
if target_nj and target_nj != "全部" and nj != target_nj:
continue
books_to_download.append(b)
total_count = len(books_to_download)
if total_count == 0:
print("[-] 未查找到符合条件的教材!请检查学段或年级名称是否正确。")
return
base_out = os.path.abspath(output_dir) if output_dir else os.path.join(get_base_dir(), "downloads")
print(f"\n📂 基础保存目录: {base_out}")
print(f"📦 待处理教材数量: {total_count} 本")
print(f"🚀 并发下载线程数: {max_workers} 个线程")
print(f"🌲 归档目录结构: {base_out}/<学段>/<年级>/<教材名>.pdf")
print(f"⚡ 断点续传机制: 开启 (已存在 PDF 自动秒跳过)")
print(f"🧹 切片缓存清理: {'开启 (每本合成后自动删除临时图片)' if clean_temp else '关闭'}")
print("-" * 70 + "\n")
downloader = PepDownloader(headless=True, output_dir=base_out)
# 统计锁
lock = threading.Lock()
processed_count = 0
success_count = 0
skipped_count = 0
failed_count = 0
start_time = time.time()
def process_single_book(item):
nonlocal processed_count, success_count, skipped_count, failed_count
idx, b = item
book_id = b.get("id")
raw_title = b.get("title", f"book_{book_id}")
# 优先采用用户习惯的学段规范大类
xdtype = b.get("xdtype", "").strip()
xd = normalize_xd(b.get("xd", "其他学段"))
if "六三" in xdtype:
stage_dir_name = "义务教育(六三学制)"
elif "五四" in xdtype or "五·四" in xdtype:
stage_dir_name = "义务教育(五四学制)"
elif "高中" in xdtype or "高中" in xd:
stage_dir_name = "高中"
else:
stage_dir_name = xd
nj = b.get("nj", "通用").strip() or "通用"
safe_xd = sanitize_filename(stage_dir_name)
safe_nj = sanitize_filename(nj)
sub_dir = os.path.join(safe_xd, safe_nj)
safe_title = sanitize_filename(raw_title)
target_dir = os.path.join(base_out, sub_dir)
expected_pdf = os.path.join(target_dir, f"{safe_title}.pdf")
# 检查是否已存在完整 PDF
if os.path.exists(expected_pdf) and os.path.getsize(expected_pdf) > 50000:
with lock:
processed_count += 1
skipped_count += 1
percent = (processed_count / total_count) * 100
size_kb = os.path.getsize(expected_pdf) // 1024
print(f"[{processed_count}/{total_count}] ({percent:5.1f}%) [✔ 已存在跳过] [{safe_xd}/{safe_nj}] 《{safe_title}》 ({size_kb} KB)")
return True
# 开始下载
try:
# 错峰启动微延时
time.sleep(random.uniform(0.1, 0.8))
pdf_path = downloader.download_book(
book_id=book_id,
custom_title=raw_title,
sub_dir=sub_dir,
skip_if_exists=True,
clean_temp=clean_temp,
quiet=True
)
with lock:
processed_count += 1
percent = (processed_count / total_count) * 100
if pdf_path and os.path.exists(pdf_path):
success_count += 1
pdf_size_mb = round(os.path.getsize(pdf_path) / (1024 * 1024), 2)
print(f"[{processed_count}/{total_count}] ({percent:5.1f}%) [🎉 下载合成完成] [{safe_xd}/{safe_nj}] 《{safe_title}》 ({pdf_size_mb} MB)")
else:
failed_count += 1
print(f"[{processed_count}/{total_count}] ({percent:5.1f}%) [❌ 下载失败] [{safe_xd}/{safe_nj}] 《{safe_title}》")
if delay > 0:
time.sleep(random.uniform(delay * 0.7, delay * 1.3))
return True
except Exception as e:
with lock:
processed_count += 1
failed_count += 1
percent = (processed_count / total_count) * 100
print(f"[{processed_count}/{total_count}] ({percent:5.1f}%) [❌ 下载异常: {e}] [{safe_xd}/{safe_nj}] 《{safe_title}》")
return False
indexed_books = list(enumerate(books_to_download, 1))
# 使用线程池并发下载
with ThreadPoolExecutor(max_workers=max_workers) as executor:
futures = [executor.submit(process_single_book, item) for item in indexed_books]
for f in as_completed(futures):
try:
f.result()
except Exception:
pass
elapsed = time.time() - start_time
elapsed_min = round(elapsed / 60, 1)
print("\n" + "=" * 70)
print("🎉 批量下载任务执行完毕!")
print(f"📊 统计汇总: 总计 {total_count} 本 | 新下载完成 {success_count} 本 | 已存在跳过 {skipped_count} 本 | 失败 {failed_count} 本")
print(f"⏱️ 总耗时: {elapsed_min} 分钟")
print(f"📁 教材归档存放于: {base_out}")
print("=" * 70)
def main():
parser = argparse.ArgumentParser(description="人教社电子教材按「学段/年级」层级目录多线程全量下载工具")
parser.add_argument("--workers", "-w", type=int, default=3, help="同时下载的并发线程数 (默认: 3)")
parser.add_argument("--output", "-o", default="./downloads", help="PDF 根输出目录 (默认: ./downloads)")
parser.add_argument("--xd", help="只下载指定学段(如:小学(六三学制)、初中(六三学制)、高中 等)")
parser.add_argument("--nj", help="只下载指定年级(如:一年级、七年级 等)")
parser.add_argument("--delay", "-d", type=float, default=0.5, help="单线程任务间休眠秒数 (默认: 0.5 秒)")
parser.add_argument("--keep-temp", action="store_true", help="保留单页切片图片缓存(默认会自动删除以节约空间)")
args = parser.parse_args()
run_download_all(
output_dir=args.output,
target_xd=args.xd,
target_nj=args.nj,
max_workers=args.workers,
delay=args.delay,
clean_temp=not args.keep_temp
)
if __name__ == "__main__":
main()
+370 -496
View File
File diff suppressed because it is too large Load Diff
+128 -33
View File
@@ -109,6 +109,36 @@ def sort_nj_key(nj: str) -> Tuple[int, int, str]:
return (1, 0, nj) return (1, 0, nj)
def map_book_xd(b: dict) -> str:
"""根据教材元数据中的 xd 与 xdtype 精确解析标准化学段"""
xd = b.get("xd", "").strip()
xdtype = b.get("xdtype", "").strip()
if "盲文" in xdtype:
return "盲校(盲文版)"
if "低视力" in xdtype:
return "盲校(低视力版)"
if "聋校" in xdtype:
return "聋校"
if "培智" in xdtype:
return "培智学校"
if "六三" in xdtype:
if xd == "初中":
return "初中(六三学制)"
return "小学(六三学制)"
if "五四" in xdtype:
if xd == "初中":
return "初中(五·四学制)"
return "小学(五四学制)"
if xd == "高中":
return "高中"
if xd == "小学":
return "小学(六三学制)"
if xd == "初中":
return "初中(六三学制)"
return normalize_xd(xd)
class PepCatalog: class PepCatalog:
"""人教社教材目录管理器""" """人教社教材目录管理器"""
BASE_URL = "https://jc.pep.com.cn/" BASE_URL = "https://jc.pep.com.cn/"
@@ -124,9 +154,8 @@ class PepCatalog:
with open(cls.LOCAL_CACHE, "r", encoding="utf-8") as f: with open(cls.LOCAL_CACHE, "r", encoding="utf-8") as f:
data = json.load(f) data = json.load(f)
if isinstance(data, list) and len(data) > 0: if isinstance(data, list) and len(data) > 0:
# 确保每条记录都包含标准化的学段
for b in data: for b in data:
b["xd"] = normalize_xd(b.get("xd", "")) b["xd"] = map_book_xd(b)
return data return data
except Exception: except Exception:
pass pass
@@ -136,22 +165,21 @@ class PepCatalog:
"Referer": cls.BASE_URL "Referer": cls.BASE_URL
} }
# 1. 下载首页定位 chunk-bfbdf2c4 JS # 1. 下载首页定位 chunk-bfbdf2c4 JS (支持兼容压缩无引号属性)
req = urllib.request.Request(cls.BASE_URL, headers=headers) req = urllib.request.Request(cls.BASE_URL, headers=headers)
with urllib.request.urlopen(req) as resp: with urllib.request.urlopen(req) as resp:
html = resp.read().decode("utf-8", errors="ignore") html = resp.read().decode("utf-8", errors="ignore")
chunk_match = re.search(r'src="(/js/chunk-bfbdf2c4\.[a-f0-9]+\.js)"', html) chunk_match = re.search(r'(/js/chunk-bfbdf2c4\.[a-f0-9]+\.js)', html)
if not chunk_match: chunk_path = chunk_match.group(1) if chunk_match else "/js/chunk-bfbdf2c4.b4dfb5d3.js"
chunk_match = re.search(r'href="(/js/chunk-bfbdf2c4\.[a-f0-9]+\.js)"', html)
chunk_path = chunk_match.group(1) if chunk_match else "/js/chunk-bfbdf2c4.3782cce3.js"
chunk_url = urllib.parse.urljoin(cls.BASE_URL, chunk_path) chunk_url = urllib.parse.urljoin(cls.BASE_URL, chunk_path)
# 2. 获取 JS 内容并提取十六进制密文 # 2. 获取 JS 内容并提取十六进制密文
js_req = urllib.request.Request(chunk_url, headers=headers) js_req = urllib.request.Request(chunk_url, headers=headers)
raw_js = urllib.request.urlopen(js_req).read() raw_js = urllib.request.urlopen(js_req).read()
m = re.search(rb'c\s*=\s*"([A-F0-9]+)"', raw_js)
if not m:
m = re.search(rb'var\s+o,\s*c\s*=\s*"([A-F0-9]+)"', raw_js) m = re.search(rb'var\s+o,\s*c\s*=\s*"([A-F0-9]+)"', raw_js)
if not m: if not m:
raise ValueError("未能从前端 JS 中匹配到教材数据密文!") raise ValueError("未能从前端 JS 中匹配到教材数据密文!")
@@ -170,9 +198,9 @@ class PepCatalog:
data_json = json.loads(plaintext.decode("utf-8")) data_json = json.loads(plaintext.decode("utf-8"))
items = data_json.get("data", []) items = data_json.get("data", [])
# 标准化学段名称 # 精确标准化学段名称
for b in items: for b in items:
b["xd"] = normalize_xd(b.get("xd", "")) b["xd"] = map_book_xd(b)
# 缓存到本地 # 缓存到本地
with open(cls.LOCAL_CACHE, "w", encoding="utf-8") as f: with open(cls.LOCAL_CACHE, "w", encoding="utf-8") as f:
@@ -284,17 +312,19 @@ class PepDownloader:
os.makedirs(self.output_dir, exist_ok=True) os.makedirs(self.output_dir, exist_ok=True)
def _solve_slider(self, page, log_cb: Optional[Callable[[str], None]] = None) -> bool: def _solve_slider(self, page, log_cb: Optional[Callable[[str], None]] = None) -> bool:
"""检测并破解阿里云 WAF 滑块""" """检测并高可靠破解阿里云 WAF 滑块(支持自动重试与防伪装刷新)"""
for check_i in range(3):
try: try:
page.wait_for_selector(".btn_slide", timeout=3000) page.wait_for_selector(".btn_slide", timeout=2500)
except Exception: except Exception:
pass pass
slider = page.query_selector(".btn_slide") slider = page.query_selector(".btn_slide")
if not slider: if not slider:
# 若无滑块,直接返回通过
return True return True
msg = "[*] 检测到阿里云 WAF 滑块验证码,正在自动滑动破解..." msg = f"[*] 检测到阿里云 WAF 滑块验证码 (第 {check_i + 1} 次尝试),正在自动滑动破解..."
if log_cb: log_cb(msg) if log_cb: log_cb(msg)
else: print(msg) else: print(msg)
@@ -303,31 +333,50 @@ class PepDownloader:
scale_box = scale.bounding_box() if scale else None scale_box = scale.bounding_box() if scale else None
if not (box and scale_box): if not (box and scale_box):
return False time.sleep(1)
continue
start_x = box["x"] + box["width"] / 2 start_x = box["x"] + box["width"] / 2
start_y = box["y"] + box["height"] / 2 start_y = box["y"] + box["height"] / 2
distance = scale_box["width"] - box["width"] + 5 distance = scale_box["width"] - box["width"] + 5
target_x = start_x + distance
page.mouse.move(start_x, start_y) page.mouse.move(start_x, start_y)
time.sleep(random.uniform(0.15, 0.3)) time.sleep(random.uniform(0.15, 0.25))
page.mouse.down() page.mouse.down()
time.sleep(0.1) time.sleep(0.05)
steps = random.randint(35, 45) # 模拟高拟真人手轨迹:带初段加速与末端微晃动
steps = random.randint(28, 38)
for i in range(1, steps + 1): for i in range(1, steps + 1):
t = i / steps t = i / steps
progress = 1 - (1 - t) * (1 - t) # 缓动函数
curr_x = start_x + distance * progress + random.uniform(-1, 1) progress = math.sin(t * (math.pi / 2))
curr_y = start_y + math.sin(t * math.pi) * 2 + random.uniform(-1, 1) curr_x = start_x + distance * progress + random.uniform(-0.8, 0.8)
curr_y = start_y + random.uniform(-1.2, 1.2)
page.mouse.move(curr_x, curr_y) page.mouse.move(curr_x, curr_y)
time.sleep(random.uniform(0.01, 0.025)) time.sleep(random.uniform(0.012, 0.022))
time.sleep(0.1) # 确保推到最右侧
page.mouse.move(start_x + distance + random.randint(2, 6), start_y)
time.sleep(0.08)
page.mouse.up() page.mouse.up()
time.sleep(2.5) time.sleep(2.5)
if page.query_selector(".btn_slide") is None:
res_msg = "[+] 滑块验证通过!"
if log_cb: log_cb(res_msg)
else: print(res_msg)
return True
else:
# 若滑块仍在,检查是否有“点击刷新”按钮
reload_btn = page.query_selector(".nc_iconfont.btn_refresh, .errloading a, .nc-lang-cnt a")
if reload_btn:
try:
reload_btn.click()
time.sleep(2)
except Exception:
pass
success = page.query_selector(".btn_slide") is None success = page.query_selector(".btn_slide") is None
res_msg = "[+] 滑块验证通过!" if success else "[-] 滑块验证未通过。" res_msg = "[+] 滑块验证通过!" if success else "[-] 滑块验证未通过。"
if log_cb: log_cb(res_msg) if log_cb: log_cb(res_msg)
@@ -337,16 +386,38 @@ class PepDownloader:
def download_book(self, def download_book(self,
book_id: str, book_id: str,
custom_title: Optional[str] = None, custom_title: Optional[str] = None,
sub_dir: Optional[str] = None,
progress_cb: Optional[Callable[[int, int, str], None]] = None, progress_cb: Optional[Callable[[int, int, str], None]] = None,
log_cb: Optional[Callable[[str], None]] = None) -> Optional[str]: log_cb: Optional[Callable[[str], None]] = None,
skip_if_exists: bool = True,
clean_temp: bool = True,
quiet: bool = False) -> Optional[str]:
""" """
下载单本教材 下载单本教材
:param book_id: 教材 ID(如 1284001101241) :param book_id: 教材 ID(如 1284001101241)
:param custom_title: 自定义书名 :param custom_title: 自定义书名
:param sub_dir: 子目录(如 "小学(六三学制)/一年级"),实现层级分类保存
:param progress_cb: 进度回调 (current_page, total_pages, status_text) :param progress_cb: 进度回调 (current_page, total_pages, status_text)
:param log_cb: 日志回调 (log_text) :param log_cb: 日志回调 (log_text)
:param skip_if_exists: 若本地已存在完整 PDF 则自动跳过
:param clean_temp: 合成 PDF 后自动删除该书的切片图片以节约磁盘空间
:param quiet: 静默模式,不打印各页下载过程,仅在关键节点或报错时提示
:return: 生成的 PDF 绝对路径 :return: 生成的 PDF 绝对路径
""" """
target_dir = os.path.join(self.output_dir, sub_dir) if sub_dir else self.output_dir
os.makedirs(target_dir, exist_ok=True)
# 检查是否已存在完整 PDF(断点续传/跳过机制)
if custom_title and skip_if_exists:
safe_title = re.sub(r'[\/:*?"<>|]', '_', custom_title).strip()
target_pdf = os.path.join(target_dir, f"{safe_title}.pdf")
if os.path.exists(target_pdf) and os.path.getsize(target_pdf) > 50000:
skip_msg = f"[✔] 本地已存在 《{safe_title}》 ({os.path.getsize(target_pdf) // 1024} KB),自动跳过。"
if log_cb: log_cb(skip_msg)
else: print(skip_msg)
if progress_cb: progress_cb(1, 1, "本地已存在,跳过")
return target_pdf
book_url = f"https://book.pep.com.cn/{book_id}/" book_url = f"https://book.pep.com.cn/{book_id}/"
init_msg = f"[*] 准备加载教材: {book_url}" init_msg = f"[*] 准备加载教材: {book_url}"
if log_cb: log_cb(init_msg) if log_cb: log_cb(init_msg)
@@ -411,6 +482,15 @@ class PepDownloader:
safe_title = re.sub(r'[\/:*?"<>|]', '_', final_title).strip() safe_title = re.sub(r'[\/:*?"<>|]', '_', final_title).strip()
total_pages = config_info.get("total", 0) total_pages = config_info.get("total", 0)
# 再次检查目标 PDF 是否存在
target_pdf = os.path.join(target_dir, f"{safe_title}.pdf")
if skip_if_exists and os.path.exists(target_pdf) and os.path.getsize(target_pdf) > 50000:
skip_msg = f"[✔] 本地已存在 《{safe_title}》 ({os.path.getsize(target_pdf) // 1024} KB),自动跳过。"
if log_cb: log_cb(skip_msg)
else: print(skip_msg)
browser.close()
return target_pdf
if total_pages == 0: if total_pages == 0:
err = f"[-] 无法读取教材总页数 (URL: {page.url})" err = f"[-] 无法读取教材总页数 (URL: {page.url})"
if log_cb: log_cb(err) if log_cb: log_cb(err)
@@ -418,11 +498,12 @@ class PepDownloader:
browser.close() browser.close()
return None return None
if not quiet:
info_msg = f"[+] 教材: 《{safe_title}》 | 总页数: {total_pages} 页" info_msg = f"[+] 教材: 《{safe_title}》 | 总页数: {total_pages} 页"
if log_cb: log_cb(info_msg) if log_cb: log_cb(info_msg)
else: print(info_msg) else: print(info_msg)
temp_dir = os.path.join(".", "temp_pages", book_id) temp_dir = os.path.join(get_base_dir(), "temp_pages", book_id)
os.makedirs(temp_dir, exist_ok=True) os.makedirs(temp_dir, exist_ok=True)
image_files = [] image_files = []
@@ -439,7 +520,7 @@ class PepDownloader:
continue continue
download_success = False download_success = False
for retry in range(3): for retry in range(4):
res = page.evaluate("""async (url) => { res = page.evaluate("""async (url) => {
try { try {
const resp = await fetch(url); const resp = await fetch(url);
@@ -467,6 +548,7 @@ class PepDownloader:
f.write(raw_data) f.write(raw_data)
image_files.append(img_path) image_files.append(img_path)
size_kb = len(raw_data) // 1024 size_kb = len(raw_data) // 1024
if not quiet:
log_text = f" -> [{page_num}/{total_pages}] 下载成功 ({size_kb} KB)" log_text = f" -> [{page_num}/{total_pages}] 下载成功 ({size_kb} KB)"
if log_cb: log_cb(log_text) if log_cb: log_cb(log_text)
else: print(log_text) else: print(log_text)
@@ -474,28 +556,34 @@ class PepDownloader:
download_success = True download_success = True
break break
# 触发 WAF 验证码拦截 # 触发 WAF 验证码或网络抖动拦截
warn_msg = f" [!] 第 {page_num} 页触发验证码,自动刷新破解..." warn_msg = f" [!] 《{safe_title}》第 {page_num} 页触发验证码/拦截 (重试 {retry + 1}/4),自动重新过盾..."
if log_cb: log_cb(warn_msg) if log_cb: log_cb(warn_msg)
else: print(warn_msg) else: print(warn_msg)
page.goto(book_url, referer="https://jc.pep.com.cn/", wait_until="networkidle")
time.sleep(1.5) # 重新刷新/导航并破解滑块
try:
page.goto(book_url, referer="https://jc.pep.com.cn/", wait_until="load")
time.sleep(2)
self._solve_slider(page, log_cb) self._solve_slider(page, log_cb)
time.sleep(2.5)
except Exception:
time.sleep(3) time.sleep(3)
if not download_success: if not download_success:
fail_msg = f" -> [{page_num}/{total_pages}] 下载失败!" fail_msg = f" -> 《{safe_title}》[{page_num}/{total_pages}] 下载失败!"
if log_cb: log_cb(fail_msg) if log_cb: log_cb(fail_msg)
else: print(fail_msg) else: print(fail_msg)
time.sleep(random.uniform(0.35, 0.65)) time.sleep(random.uniform(0.15, 0.35))
browser.close() browser.close()
if not image_files: if not image_files:
return None return None
output_pdf = os.path.join(self.output_dir, f"{safe_title}.pdf") output_pdf = target_pdf
if not quiet:
merge_msg = f"[*] 正在合成 PDF: {os.path.basename(output_pdf)} ..." merge_msg = f"[*] 正在合成 PDF: {os.path.basename(output_pdf)} ..."
if log_cb: log_cb(merge_msg) if log_cb: log_cb(merge_msg)
else: print(merge_msg) else: print(merge_msg)
@@ -518,6 +606,13 @@ class PepDownloader:
rest_images = pil_images[1:] rest_images = pil_images[1:]
first_im.save(output_pdf, "PDF", resolution=100.0, save_all=True, append_images=rest_images) first_im.save(output_pdf, "PDF", resolution=100.0, save_all=True, append_images=rest_images)
# 下载合成完毕后,自动清理单页图片切片,释放磁盘空间
if clean_temp and os.path.exists(temp_dir):
try:
shutil.rmtree(temp_dir)
except Exception:
pass
done_msg = f"[✔] PDF 生成成功: {output_pdf}" done_msg = f"[✔] PDF 生成成功: {output_pdf}"
if log_cb: log_cb(done_msg) if log_cb: log_cb(done_msg)
else: print(done_msg) else: print(done_msg)
+12 -3
View File
@@ -17,7 +17,7 @@ from fastapi import FastAPI, BackgroundTasks, WebSocket, WebSocketDisconnect
from fastapi.responses import HTMLResponse, JSONResponse from fastapi.responses import HTMLResponse, JSONResponse
from fastapi.staticfiles import StaticFiles from fastapi.staticfiles import StaticFiles
from pep_core import PepCatalog, PepDownloader, XD_ORDER, XK_ORDER_PREFIX, NJ_ORDER, get_base_dir from pep_core import PepCatalog, PepDownloader, XD_ORDER, XK_ORDER_PREFIX, NJ_ORDER, get_base_dir, normalize_xd
if hasattr(sys.stdout, "reconfigure"): if hasattr(sys.stdout, "reconfigure"):
sys.stdout.reconfigure(encoding="utf-8") sys.stdout.reconfigure(encoding="utf-8")
@@ -104,13 +104,22 @@ class TaskManager:
self.log(txt) self.log(txt)
try: try:
xd = normalize_xd(book_to_download.get("xd", "其他学段"))
nj = book_to_download.get("nj", "通用").strip() or "通用"
safe_xd = re.sub(r'[\/:*?"<>|]', '_', xd).strip()
safe_nj = re.sub(r'[\/:*?"<>|]', '_', nj).strip()
sub_dir = os.path.join(safe_xd, safe_nj)
self.log(f"==================================================") self.log(f"==================================================")
self.log(f"[*] 开始下载教材: 《{book_to_download.get('title')}》") self.log(f"[*] 开始下载教材: [{safe_xd}/{safe_nj}] 《{book_to_download.get('title')}》")
downloader.download_book( downloader.download_book(
book_id=book_to_download["id"], book_id=book_to_download["id"],
custom_title=book_to_download.get("title"), custom_title=book_to_download.get("title"),
sub_dir=sub_dir,
progress_cb=progress_callback, progress_cb=progress_callback,
log_cb=log_callback log_cb=log_callback,
skip_if_exists=True,
clean_temp=True
) )
except Exception as e: except Exception as e:
self.log(f"[-] 下载异常: {e}") self.log(f"[-] 下载异常: {e}")