ops: 新增 course-toolkit(YouTube 课程下载+转写+成册工具,含踩坑记录)
This commit is contained in:
@@ -0,0 +1,55 @@
|
||||
# course-toolkit — YouTube 课程「下载 + 转写 + 成册」工具包
|
||||
|
||||
把某个 YouTube 课程/播放列表,批量变成 **本地视频 + 英文文稿 + 中文导读 Markdown**。
|
||||
|
||||
## 组成
|
||||
| 文件 | 作用 |
|
||||
|---|---|
|
||||
| `fetch_course.sh` | 主入口:枚举播放列表 → 并行下载+转写 → 清洗文稿 |
|
||||
| `worker.sh` | 单个视频的下载/抽音频/转写(被主脚本并行调用) |
|
||||
| `course_util.py` | 子命令:`manifest` / `clean` / `transcribe` / `build` |
|
||||
|
||||
## 快速开始
|
||||
```bash
|
||||
# 1) 下载 + 转写(默认英文,4 路并行)
|
||||
./fetch_course.sh "https://www.youtube.com/playlist?list=XXXX" ~/course/我的课程
|
||||
# 中文课程: MODEL=small LANG=zh ./fetch_course.sh <url> <outdir>
|
||||
|
||||
# 2) 生成中文导读:在 <outdir>/guides/ 下放 01.txt 02.txt ...(可直接让 AI 读 clean/NN.txt 生成)
|
||||
# 然后组装成册:
|
||||
python3 course_util.py build ~/course/我的课程
|
||||
```
|
||||
|
||||
产物结构:
|
||||
```
|
||||
<outdir>/
|
||||
├── manifest.txt # NN|id|slug|title
|
||||
├── videos/NN-slug.mp4 # 本地视频(默认 360p)
|
||||
├── transcripts/ # 带时间轴文稿
|
||||
├── clean/NN.txt # 去时间轴文稿(供检索/摘要)
|
||||
├── guides/NN.txt # 中文导读(可选,人工/AI 写)
|
||||
├── NN-slug.md # 成册笔记(导读 + 全文)
|
||||
├── README.md # 目录索引
|
||||
└── build.log # 运行日志
|
||||
```
|
||||
|
||||
## 环境变量
|
||||
| 变量 | 默认 | 说明 |
|
||||
|---|---|---|
|
||||
| `PROXY` | `socks5://192.168.27.96:1080` | 访问 YouTube 的代理 |
|
||||
| `MODEL` | `base.en` | faster-whisper 模型(英文 base.en/small.en;中文 small)|
|
||||
| `LANG` | `en` | 转写语言 |
|
||||
| `JOBS` | `4` | 并行数 |
|
||||
| `FMT` | `18` | yt-dlp 格式(见下)|
|
||||
|
||||
依赖:`yt-dlp`、`ffmpeg`、`faster-whisper` + `python3`。
|
||||
|
||||
## 关键经验(踩过的坑)
|
||||
1. **画质问题**:YouTube 未登录下载被限,实测只有 `--extractor-args "youtube:player_client=android"` 能下 → **上限 360p**(格式 18,渐进式、含音轨)。`android_vr`/`tv_embedded` 会列出 1080p 但下载 **403**;要高清需 PO token / 登录 cookie。
|
||||
2. **官方字幕不稳**:常遇 **429 限流**或该视频**根本没有字幕** → 故用**本地 ASR**替代。
|
||||
3. **模型下载**:走 `hf-mirror`,并必须 `HF_HUB_DISABLE_XET=1`(否则 xet CAS 401)。
|
||||
4. **速度**:faster-whisper `base.en` + `beam_size=1` + **单线程多进程并行**,实测 ~28x 实时(4 并行);19 集≈12 分钟。
|
||||
5. **成册**:每集 md = 中文导读 + 英文全文;文稿为自动转写,**未人工校对**。
|
||||
6. ⚠️ 版权:视频版权归原作者,建议仅作**个人学习**归档,勿公开分发。
|
||||
|
||||
_(本工具包由小五沉淀,首个用户:《Full Algorithmic Trading Using Python》19 集)_
|
||||
@@ -0,0 +1,102 @@
|
||||
#!/usr/bin/env python3
|
||||
"""course-toolkit 辅助工具。
|
||||
|
||||
子命令:
|
||||
manifest <rawfile> 将 index|id|title 转为 NN|id|slug|title
|
||||
clean <outdir> transcripts/*.txt -> clean/*.txt(去时间轴、并段)
|
||||
transcribe <wav> <out> 本地 ASR 转写(env: MODEL, LANG, THREADS)
|
||||
build <outdir> 组装 md + README(guides/<NN>.txt 可选作中文导读)
|
||||
"""
|
||||
import os, re, sys
|
||||
|
||||
def slugify(s):
|
||||
s = s.lower()
|
||||
s = re.sub(r'[^a-z0-9]+', '-', s)
|
||||
return re.sub(r'-+', '-', s).strip('-')[:60] or 'video'
|
||||
|
||||
def cmd_manifest(rawfile):
|
||||
out = []
|
||||
for line in open(rawfile, encoding='utf-8'):
|
||||
line = line.strip()
|
||||
if not line: continue
|
||||
parts = line.split('|')
|
||||
if len(parts) < 3: continue
|
||||
idx, vid, title = parts[0], parts[1], '|'.join(parts[2:])
|
||||
try: nn = f"{int(idx):02d}"
|
||||
except Exception: nn = slugify(idx)[:2]
|
||||
out.append(f"{nn}|{vid}|{slugify(title)}|{title}")
|
||||
sys.stdout.write("\n".join(out) + "\n")
|
||||
|
||||
def cmd_clean(outdir):
|
||||
tdir = os.path.join(outdir, 'transcripts'); cdir = os.path.join(outdir, 'clean')
|
||||
os.makedirs(cdir, exist_ok=True)
|
||||
for f in sorted(os.listdir(tdir)):
|
||||
if not f.endswith('.txt'): continue
|
||||
nn = f.split('-')[0]
|
||||
lines = [l for l in open(os.path.join(tdir, f), encoding='utf-8').read().splitlines() if l.strip()]
|
||||
txt = " ".join(re.sub(r'^\[\s*[\d.]+\s*-\s*[\d.]+\s*\]\s*', '', l) for l in lines)
|
||||
open(os.path.join(cdir, f"{nn}.txt"), 'w', encoding='utf-8').write(txt)
|
||||
print(f"[clean] {len(os.listdir(cdir))} files -> {cdir}")
|
||||
|
||||
def cmd_transcribe(wav, out):
|
||||
os.environ.setdefault("HF_ENDPOINT", "https://hf-mirror.com")
|
||||
os.environ.setdefault("HF_HUB_DISABLE_XET", "1")
|
||||
from faster_whisper import WhisperModel
|
||||
model = os.environ.get("MODEL", "base.en")
|
||||
lang = os.environ.get("LANG", "en")
|
||||
threads = int(os.environ.get("THREADS", "1"))
|
||||
m = WhisperModel(model, device="cpu", compute_type="int8", cpu_threads=threads)
|
||||
segs, info = m.transcribe(wav, language=lang, beam_size=1, vad_filter=True)
|
||||
lines = [f"[{s.start:7.1f}-{s.end:7.1f}] {s.text.strip()}" for s in segs]
|
||||
open(out, 'w', encoding='utf-8').write("\n".join(lines))
|
||||
print(f"[transcribe] {out} segs={len(lines)} dur={info.duration:.0f}")
|
||||
|
||||
def cmd_build(outdir):
|
||||
import subprocess, json
|
||||
man = [l.split('|') for l in open(os.path.join(outdir, 'manifest.txt'), encoding='utf-8').read().splitlines() if l.strip()]
|
||||
gdir = os.path.join(outdir, 'guides')
|
||||
rows = []
|
||||
for p in man:
|
||||
NN, ID, SLUG = p[0], p[1], p[2]
|
||||
title = p[3] if len(p) > 3 else SLUG
|
||||
vid = f"videos/{NN}-{SLUG}.mp4"
|
||||
txt = open(os.path.join(outdir, 'clean', f'{NN}.txt'), encoding='utf-8').read().strip()
|
||||
gpath = os.path.join(gdir, f'{NN}.txt')
|
||||
guide = open(gpath, encoding='utf-8').read().strip() if os.path.exists(gpath) else "> TODO:待补充中文导读。"
|
||||
try:
|
||||
dur = float(subprocess.check_output(["ffprobe", "-v", "error", "-show_entries", "format=duration", "-of", "csv=p=0", os.path.join(outdir, vid)]).decode())
|
||||
mm, ss = int(dur // 60), int(dur % 60)
|
||||
except Exception:
|
||||
mm = ss = 0
|
||||
md = f"""# {NN} · {title}
|
||||
|
||||
- **原始视频**:https://youtu.be/{ID}
|
||||
- **时长**:{mm} 分 {ss} 秒
|
||||
- **本地视频**:[{vid}]({vid})
|
||||
|
||||
## 🎯 本集要点(中文导读)
|
||||
|
||||
{guide}
|
||||
|
||||
## 📝 完整文稿(自动转写)
|
||||
|
||||
> 由本地 ASR 转写,未人工校对,供检索/精读使用。
|
||||
|
||||
{txt}
|
||||
"""
|
||||
open(os.path.join(outdir, f'{NN}-{SLUG}.md'), 'w', encoding='utf-8').write(md)
|
||||
rows.append((NN, title, mm, ss, SLUG))
|
||||
idx = "\n".join(f"| {n} | {t} | {m}:{s:02d} | [📄 笔记]({n}-{sl}.md) · [🎬 视频](videos/{n}-{sl}.mp4) |" for n, t, m, s, sl in rows)
|
||||
open(os.path.join(outdir, 'README.md'), 'w', encoding='utf-8').write(
|
||||
"# 课程教程\n\n> 由 course-toolkit 生成。\n\n| 集 | 标题 | 时长 | 链接 |\n|---|---|---|---|\n" + idx + "\n")
|
||||
print(f"[build] {len(rows)} 篇 + README -> {outdir}")
|
||||
|
||||
if __name__ == '__main__':
|
||||
if len(sys.argv) < 2:
|
||||
print(__doc__); sys.exit(1)
|
||||
cmd = sys.argv[1]
|
||||
if cmd == 'manifest': cmd_manifest(sys.argv[2])
|
||||
elif cmd == 'clean': cmd_clean(sys.argv[2])
|
||||
elif cmd == 'transcribe': cmd_transcribe(sys.argv[2], sys.argv[3])
|
||||
elif cmd == 'build': cmd_build(sys.argv[2])
|
||||
else: print(__doc__); sys.exit(1)
|
||||
@@ -0,0 +1,34 @@
|
||||
#!/bin/bash
|
||||
# fetch_course.sh — 下载 YouTube 课程视频 + 本地转写 + 组装教程
|
||||
#
|
||||
# 用法:
|
||||
# ./fetch_course.sh <playlist_url或单个视频url> <输出目录>
|
||||
#
|
||||
# 环境变量(可选):
|
||||
# PROXY 代理,默认 socks5://192.168.27.96:1080
|
||||
# MODEL whisper 模型,默认 base.en(英文用 base.en/small.en,中文用 small)
|
||||
# LANG 语言,默认 en(中文用 zh)
|
||||
# JOBS 并行数,默认 4
|
||||
# FMT yt-dlp 格式,默认 18(360p 渐进式,含音轨;见 README 画质说明)
|
||||
set -uo pipefail
|
||||
URL="${1:?用法: ./fetch_course.sh <url> <输出目录>}"
|
||||
OUT="${2:?缺少输出目录}"
|
||||
PROXY="${PROXY:-socks5://192.168.27.96:1080}"
|
||||
MODEL="${MODEL:-base.en}"; LG="${LANG:-en}"; JOBS="${JOBS:-4}"; FMT="${FMT:-18}"
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
mkdir -p "$OUT/videos" "$OUT/transcripts" "$OUT/clean"
|
||||
|
||||
if [ ! -s "$OUT/manifest.txt" ]; then
|
||||
echo "[*] 枚举视频列表..."
|
||||
yt-dlp --proxy "$PROXY" --no-warnings --flat-playlist \
|
||||
--print "%(playlist_index)s|%(id)s|%(title)s" "$URL" > "$OUT/_raw.txt" || { echo "枚举失败"; exit 1; }
|
||||
python3 "$HERE/course_util.py" manifest "$OUT/_raw.txt" > "$OUT/manifest.txt"
|
||||
fi
|
||||
echo "[*] 待处理视频: $(grep -c . "$OUT/manifest.txt") 个 | 模型=$MODEL 并行=$JOBS"
|
||||
|
||||
export PROXY MODEL LANG="$LG" OUT HERE FMT
|
||||
grep -v '^[[:space:]]*$' "$OUT/manifest.txt" | xargs -P "$JOBS" -I{} "$HERE/worker.sh" "{}"
|
||||
|
||||
python3 "$HERE/course_util.py" clean "$OUT"
|
||||
echo "[*] 全部完成: $OUT"
|
||||
echo "[*] 下一步: 在 $OUT/guides/ 放中文导读(<NN>.txt)后运行 course_util.py build $OUT"
|
||||
@@ -0,0 +1,23 @@
|
||||
#!/bin/bash
|
||||
# worker.sh — 处理单个视频:下载 → 抽音频 → 转写(由 fetch_course.sh 调用)
|
||||
set -uo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
line="$1"
|
||||
NN=$(echo "$line"|cut -d'|' -f1); ID=$(echo "$line"|cut -d'|' -f2); SLUG=$(echo "$line"|cut -d'|' -f3)
|
||||
VID="$OUT/videos/${NN}-${SLUG}.mp4"; TXT="$OUT/transcripts/${NN}-${SLUG}.${LANG}.txt"
|
||||
LOG="$OUT/build.log"
|
||||
log(){ echo "[$(date '+%F %T')] [$NN] $*" >> "$LOG"; }
|
||||
[ -f "$TXT" ] && { log "skip(done)"; exit 0; }
|
||||
if [ ! -f "$VID" ]; then
|
||||
for a in 1 2 3; do
|
||||
yt-dlp --proxy "$PROXY" --no-warnings --no-part \
|
||||
--extractor-args "youtube:player_client=android" -f "$FMT" \
|
||||
-o "$VID" "https://www.youtube.com/watch?v=$ID" >>"$LOG" 2>&1 && break
|
||||
log "retry-dl $a"; sleep 8
|
||||
done
|
||||
fi
|
||||
[ -f "$VID" ] || { log "FAIL download"; exit 1; }
|
||||
ffmpeg -y -hide_banner -loglevel error -i "$VID" -ar 16000 -ac 1 -c:a pcm_s16le "$OUT/_${NN}.wav" 2>>"$LOG" || { log "FAIL audio"; exit 1; }
|
||||
MODEL="$MODEL" LANG="$LANG" THREADS=1 python3 "$HERE/course_util.py" transcribe "$OUT/_${NN}.wav" "$TXT" >>"$LOG" 2>&1 || { log "FAIL transcribe"; exit 1; }
|
||||
rm -f "$OUT/_${NN}.wav"
|
||||
log "ok $(du -h "$VID"|cut -f1)"
|
||||
Reference in New Issue
Block a user