Compare commits
12
Commits
5d9352ccf3
..
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b86c6929e4 | ||
|
|
0d8bcc5d13 | ||
|
|
25b0aacf52 | ||
|
|
b6efec098f | ||
|
|
cb211da473 | ||
|
|
5629d57ea7 | ||
|
|
9a43ecfedf | ||
|
|
4ecab4c47c | ||
|
|
84912ef99c | ||
|
|
e07e9426c7 | ||
|
|
71444a8172 | ||
|
|
51079bd8dc |
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,37 @@
|
||||
# Full Algorithmic Trading Using Python — 课程教程
|
||||
|
||||
> 本目录是对 YouTube 频道 **TradeOptionsWithMe** 的《Full Algorithmic Trading Using Python》系列的**学习笔记 + 视频归档**。
|
||||
> 播放列表:https://youtube.com/playlist?list=PLtqRgJ_TIq8Y6YG8-G-ETIFW_36mvxMLad
|
||||
|
||||
- **集数**:19 集 | **总时长**:约 6 小时 45 分
|
||||
- **每集内容**:📌 中文要点导读(提炼) + 📝 英文完整文稿(本地 ASR 转写) + 🎬 本地视频(360p)
|
||||
- **提醒**:文稿由自动语音识别生成,**未人工校对**;仅供个人学习检索,版权归原作者所有。视频画质上限 360p(YouTube 未登录下载限制)。
|
||||
|
||||
## 目录
|
||||
|
||||
| 集 | 标题 | 时长 | 链接 |
|
||||
|---|---|---|---|
|
||||
| 01 | Introduction(课程总览) | 7:19 | [📄 笔记](01-introduction.md) · [🎬 视频](videos/01-introduction.mp4) |
|
||||
| 02 | How Trading Algorithms Work(算法原理与开发流程) | 10:03 | [📄 笔记](02-episode-2.md) · [🎬 视频](videos/02-episode-2.mp4) |
|
||||
| 03 | Key Concepts(关键概念:时间处理·骨架·Symbol) | 10:05 | [📄 笔记](03-key-concepts.md) · [🎬 视频](videos/03-key-concepts.mp4) |
|
||||
| 04 | Handling Data(处理数据 · 第一个算法) | 18:01 | [📄 笔记](04-handling-data.md) · [🎬 视频](videos/04-handling-data.mp4) |
|
||||
| 05 | Trading & Orders(下单、订单管理与调试) | 28:46 | [📄 笔记](05-trading-and-orders.md) · [🎬 视频](videos/05-trading-and-orders.mp4) |
|
||||
| 06 | Indicators & Historical Data(指标与历史数据) | 28:32 | [📄 笔记](06-indicators-history.md) · [🎬 视频](videos/06-indicators-history.mp4) |
|
||||
| 07 | Consolidators & Rolling Windows(合并器、滚动窗口、事件调度) | 22:40 | [📄 笔记](07-consolidators-rolling-windows.md) · [🎬 视频](videos/07-consolidators-rolling-windows.mp4) |
|
||||
| 08 | Dynamic Universes(动态证券池与基本面选股) | 25:30 | [📄 笔记](08-dynamic-universes.md) · [🎬 视频](videos/08-dynamic-universes.mp4) |
|
||||
| 09 | Twitter Trading Bot(自定义数据 · 推文情绪) | 26:52 | [📄 笔记](09-twitter-trading-bot.md) · [🎬 视频](videos/09-twitter-trading-bot.mp4) |
|
||||
| 10 | Backtesting & Performance Analysis(回测与绩效评估) | 25:45 | [📄 笔记](10-backtesting-performance.md) · [🎬 视频](videos/10-backtesting-performance.mp4) |
|
||||
| 11 | Forex Trading(外汇 · 均值回归) | 13:27 | [📄 笔记](11-forex-trading.md) · [🎬 视频](videos/11-forex-trading.mp4) |
|
||||
| 12 | Options Trading(期权入门) | 15:28 | [📄 笔记](12-options-trading.md) · [🎬 视频](videos/12-options-trading.mp4) |
|
||||
| 13 | Options Code-Along(保护性看跌期权实战) | 26:06 | [📄 笔记](13-options-code-along.md) · [🎬 视频](videos/13-options-code-along.mp4) |
|
||||
| 14 | Crypto Trading Bots(加密货币 · RSI 动量) | 15:51 | [📄 笔记](14-crypto-trading-bots.md) · [🎬 视频](videos/14-crypto-trading-bots.mp4) |
|
||||
| 15 | The Algorithm Framework(算法框架) | 34:14 | [📄 笔记](15-algorithm-framework.md) · [🎬 视频](videos/15-algorithm-framework.mp4) |
|
||||
| 16 | Data-Driven Research(研究环境 · QuantBook) | 21:52 | [📄 笔记](16-data-driven-research.md) · [🎬 视频](videos/16-data-driven-research.mp4) |
|
||||
| 17 | Bitcoin ML Bot(用神经网络做比特币) | 31:07 | [📄 笔记](17-bitcoin-ml-bot.md) · [🎬 视频](videos/17-bitcoin-ml-bot.mp4) |
|
||||
| 18 | How to Live Trade(实盘部署与 Paper Trading) | 26:03 | [📄 笔记](18-live-trading.md) · [🎬 视频](videos/18-live-trading.mp4) |
|
||||
| 19 | 50-Line Trading Bot(50 行代码 · 标普成分股调整) | 17:29 | [📄 笔记](19-trading-bot-50-lines.md) · [🎬 视频](videos/19-trading-bot-50-lines.mp4) |
|
||||
|
||||
## 关于本归档
|
||||
- 视频:经 yt-dlp 下载归档(360p),讲稿:faster-whisper 本地转写,导读:小五整理。
|
||||
- 系列主线为 **Python + QuantConnect(Lean 引擎)** 开发算法交易:从概念 → 下单/指标/数据 → 各类资产(股/汇/期权/币)→ 框架/研究/ML → 实盘。
|
||||
- 生成日期:2026-09-12。
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
BIN
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,20 @@
|
||||
# course — 量化学习课程与资料
|
||||
|
||||
本目录按「课程 + 资料 + 练习」分类归档。
|
||||
|
||||
## 目录结构
|
||||
|
||||
```
|
||||
course/
|
||||
├── Full Algorithmic Trading Using Python/ # 视频课程(19集,笔记+视频存档)
|
||||
├── 学习资料/ # 学习资料总表 / 视频候选清单
|
||||
└── 动手练习/ # 可运行的 A 股量化练习代码
|
||||
```
|
||||
|
||||
| 分类 | 说明 |
|
||||
|---|---|
|
||||
| `Full Algorithmic Trading Using Python/` | YouTube「TradeOptionsWithMe」系列全套:19 集中文要点笔记 + 英文转写 + 本地视频 |
|
||||
| `学习资料/` | `资料总表.md`(书单/平台/框架/社区)、`视频候选清单.md`(B站待下载候选) |
|
||||
| `动手练习/` | 01 pandas热身 · 02 akshare取数 · 03 双均线回测 · 04 动量轮动 · 05 多因子;含 `lib_quant.py`、`run.sh` |
|
||||
|
||||
> 整理:小五 | 2026-09-28
|
||||
@@ -0,0 +1,26 @@
|
||||
"""01 | pandas 热身:收益率、波动、净值曲线
|
||||
目标:搞懂量化里最基础的三个东西——收益、风险、净值。
|
||||
"""
|
||||
import pandas as pd, numpy as np
|
||||
import matplotlib
|
||||
matplotlib.use("Agg")
|
||||
import matplotlib.pyplot as plt
|
||||
from lib_quant import get_daily, add_returns, net_value, stats
|
||||
|
||||
df = get_daily("600519", "20200101", "20241231") # 贵州茅台
|
||||
df = add_returns(df)
|
||||
nv = net_value(df["ret"])
|
||||
print("==", "600519 贵州茅台", "==")
|
||||
print(df[["日期", "收盘", "ret"]].tail(3).to_string(index=False))
|
||||
print("绩效:", stats(df["ret"]))
|
||||
|
||||
plt.figure(figsize=(9, 4))
|
||||
plt.plot(df["日期"], nv)
|
||||
plt.title("Net value (buy & hold)")
|
||||
plt.xlabel("date"); plt.ylabel("nav")
|
||||
plt.tight_layout(); plt.savefig("../data/01_net_value.png", dpi=110)
|
||||
print("图已存 data/01_net_value.png")
|
||||
|
||||
# 思考题:
|
||||
# 1) 为什么不看"价格涨了多少",而要看"收益率"和"净值"?
|
||||
# 2) 日收益率为 0 的日子占多少?(df['ret'].eq(0).mean())
|
||||
@@ -0,0 +1,19 @@
|
||||
"""02 | 用 akshare 拉数据并落地缓存
|
||||
目标:掌握免费 A股数据源用法——这是量化的一切起点。
|
||||
"""
|
||||
from lib_quant import get_daily, CACHE
|
||||
import os
|
||||
|
||||
for code, name in [("000001", "平安银行"), ("600519", "贵州茅台"), ("300750", "宁德时代")]:
|
||||
df = get_daily(code, "20230101", "20241231")
|
||||
d0 = str(df["日期"].iloc[0])[:10]
|
||||
d1 = str(df["日期"].iloc[-1])[:10]
|
||||
print(f"{code} {name}: {len(df)} 行 | {d0} ~ {d1}")
|
||||
print(" 列:", list(df.columns))
|
||||
|
||||
print("\n缓存目录:", os.path.abspath(CACHE))
|
||||
print("文件:", sorted(os.listdir(CACHE)))
|
||||
|
||||
# 思考题:
|
||||
# 1) 前复权(qfq)/后复权(hfq)/不复权 区别?算收益率该用哪个?(改 adjust 参数对比)
|
||||
# 2) 每日数据延迟多久?实盘能用吗?(提示: akshare 文档 + 交易所规则)
|
||||
@@ -0,0 +1,33 @@
|
||||
"""03 | 双均线策略 + 回测
|
||||
目标:写出人生第一个完整策略,并做归因。
|
||||
策略:5日均线上穿20日均线买入,下穿卖出。
|
||||
"""
|
||||
import pandas as pd, numpy as np
|
||||
import matplotlib; matplotlib.use("Agg")
|
||||
import matplotlib.pyplot as plt
|
||||
from lib_quant import get_daily, add_returns, net_value, stats
|
||||
|
||||
FAST, SLOW = 5, 20
|
||||
df = get_daily("600519", "20200101", "20241231")
|
||||
df = add_returns(df)
|
||||
df["ma_fast"] = df["收盘"].rolling(FAST).mean()
|
||||
df["ma_slow"] = df["收盘"].rolling(SLOW).mean()
|
||||
df["signal"] = (df["ma_fast"] > df["ma_slow"]).astype(int) # 1=持有 0=空仓
|
||||
df["pos"] = df["signal"].shift(1).fillna(0) # 次日开盘才算,防未来函数
|
||||
df["strat_ret"] = df["pos"] * df["ret"]
|
||||
|
||||
nv_bh, nv_st = net_value(df["ret"]), net_value(df["strat_ret"])
|
||||
print("== 买入持有 ==", stats(df["ret"]))
|
||||
print("== 双均线 ==", stats(df["strat_ret"]))
|
||||
|
||||
plt.figure(figsize=(10, 4))
|
||||
plt.plot(df["日期"], nv_bh, label="BuyHold")
|
||||
plt.plot(df["日期"], nv_st, label=f"MA{FAST}/{SLOW}")
|
||||
plt.legend(); plt.title("Strategy vs Buy&Hold"); plt.tight_layout()
|
||||
plt.savefig("../data/03_dual_ma.png", dpi=110)
|
||||
print("图已存 data/03_dual_ma.png")
|
||||
|
||||
# 思考题:
|
||||
# 1) 为什么 pos 要 shift(1)?不 shift 会怎样?(把 shift(1) 去掉对比——这叫"未来函数")
|
||||
# 2) 改 FAST/SLOW 参数,结果稳吗?换 (10,60) 试试——这叫参数敏感性
|
||||
# 3) 这策略赢过买入持有了吗?为什么?
|
||||
@@ -0,0 +1,36 @@
|
||||
"""04 | 动量轮动策略
|
||||
目标:从"单票择时"升级到"多标的轮动"。
|
||||
策略:每月末看过去 N 日涨幅,持有最强的一只。
|
||||
"""
|
||||
import pandas as pd, numpy as np
|
||||
from lib_quant import get_daily, net_value, stats
|
||||
|
||||
POOL = {"600519": "贵州茅台", "000001": "平安银行", "300750": "宁德时代",
|
||||
"601318": "中国平安", "000858": "五粮液"}
|
||||
LOOKBACK, FREQ = 20, 20 # 回看20日,每20个交易日调仓
|
||||
|
||||
px = {}
|
||||
for c in POOL:
|
||||
d = get_daily(c, "20200101", "20241231")[["日期", "收盘"]].set_index("日期")
|
||||
px[c] = d["收盘"]
|
||||
prices = pd.DataFrame(px).dropna()
|
||||
mom = prices.pct_change(LOOKBACK)
|
||||
|
||||
rets = []
|
||||
cur = None
|
||||
for i in range(LOOKBACK, len(prices) - 1):
|
||||
if (i - LOOKBACK) % FREQ == 0: # 调仓日
|
||||
cur = mom.iloc[i].idxmax() # 动量最强
|
||||
if cur is not None:
|
||||
r = prices[cur].iloc[i + 1] / prices[cur].iloc[i] - 1
|
||||
rets.append(r)
|
||||
else:
|
||||
rets.append(0.0)
|
||||
|
||||
s = pd.Series(rets)
|
||||
print("== 动量轮动 ==", stats(s))
|
||||
print("== 等权买入持有 ==", stats(prices.pct_change().mean(axis=1).iloc[LOOKBACK:]))
|
||||
|
||||
# 思考题:
|
||||
# 1) 换 LOOKBACK (10/60/120) 看结果——动量在A股有效吗?
|
||||
# 2) 为什么不持有全部、按动量加权?(提示: 集中 vs 分散)
|
||||
@@ -0,0 +1,37 @@
|
||||
"""05 | 多因子选股(横截面)
|
||||
目标:体验"因子"最朴素的样子——按指标排序选股。
|
||||
数据:akshare 全A快照(含市盈率/市净率等)。
|
||||
"""
|
||||
import time
|
||||
import akshare as ak, pandas as pd
|
||||
|
||||
def fetch_spot(retries=6):
|
||||
last = None
|
||||
for i in range(retries):
|
||||
try:
|
||||
return ak.stock_zh_a_spot_em()
|
||||
except Exception as e:
|
||||
last = e; time.sleep(3 + i * 3)
|
||||
raise last
|
||||
|
||||
try:
|
||||
spot = fetch_spot()
|
||||
print("全A快照:", spot.shape)
|
||||
df = spot.copy()
|
||||
pe = pd.to_numeric(df.get("市盈率-动态"), errors="coerce")
|
||||
df = df[pe.notna()].copy()
|
||||
df["EP"] = 1 / pe[pe.notna()] # 盈利收益率(1/PE)
|
||||
df["EP_rank"] = df["EP"].rank(pct=True)
|
||||
top = df.sort_values("EP_rank", ascending=False).head(20)[
|
||||
["代码", "名称", "最新价", "市盈率-动态", "EP_rank"]]
|
||||
print("\n== 按盈利收益率(EP)选出的前20(低估值)==")
|
||||
print(top.to_string(index=False))
|
||||
top.to_csv("../data/05_top20_EP.csv", index=False, encoding="utf-8-sig")
|
||||
print("已存 data/05_top20_EP.csv")
|
||||
except Exception as e:
|
||||
print("取数失败(网络/接口变动):", repr(e)[:200])
|
||||
print("可稍后重跑;或先跑 01-04(已缓存,离线可跑)")
|
||||
|
||||
# 思考题:
|
||||
# 1) 单因子排序选股有什么坑?(行业偏差/市值偏差/未来函数)
|
||||
# 2) 怎么把2个因子合成打分?(标准化后加权,试试 z-score)
|
||||
@@ -0,0 +1,23 @@
|
||||
# 量化动手练习(用真实 A 股数据)
|
||||
|
||||
> 目标:把"看书看不进去"的瓶颈,换成**能跑出结果**的闭环练习。
|
||||
> 环境:Python 3.11 + pandas/akshare/backtrader/matplotlib(已装在 `../pylibs`)
|
||||
> 运行:`bash run.sh`(全部)或 `bash run.sh 03`(单课)
|
||||
|
||||
## 怎么用
|
||||
每一课 = 目标 + 代码 + 思考题。先跑通,再改参数看结果变化,最后回答思考题。卡住直接问小五。
|
||||
|
||||
## 课程表
|
||||
| 课 | 文件 | 主题 | 你会学到 |
|
||||
|---|---|---|---|
|
||||
| 01 | 01_pandas_热身.py | 收益率/波动/净值 | 量化最基础的语言 |
|
||||
| 02 | 02_拉数据_akshare.py | 用 akshare 拉 A股数据 | 免费数据源怎么用 |
|
||||
| 03 | 03_双均线回测.py | 双均线策略+回测 | 一个完整策略长什么样 |
|
||||
| 04 | 04_动量轮动.py | 动量轮动策略 | 多标的轮动调仓 |
|
||||
| 05 | 05_多因子选股.py | 多因子选股 | 从横截面数据里选股 |
|
||||
|
||||
## 数据
|
||||
- 默认用 akshare 在线拉(国内可达),首次拉完自动缓存到 `../data/cache/*.csv`
|
||||
|
||||
## 学习闭环
|
||||
跑通 → 改参数 → 想为什么 → 写结论 → 告诉小五批改/加题
|
||||
@@ -0,0 +1,48 @@
|
||||
"""量化练习公共库:数据获取(带缓存+重试) + 常用指标 + 回测指标。"""
|
||||
import os, time, pandas as pd, numpy as np
|
||||
|
||||
CACHE = os.path.join(os.path.dirname(__file__), "..", "data", "cache")
|
||||
os.makedirs(CACHE, exist_ok=True)
|
||||
|
||||
def _fetch(symbol, start, end, adjust, retries=6, wait=2.0):
|
||||
import akshare as ak
|
||||
last = None
|
||||
for i in range(retries):
|
||||
try:
|
||||
return ak.stock_zh_a_hist(symbol=symbol, period="daily",
|
||||
start_date=start, end_date=end, adjust=adjust)
|
||||
except Exception as e: # 网络间歇性抽风,重试
|
||||
last = e
|
||||
time.sleep(wait * (i + 1))
|
||||
raise last
|
||||
|
||||
def get_daily(symbol="000001", start="20200101", end="20241231", adjust="qfq"):
|
||||
"""拉单只A股日线(前复权),带本地缓存与自动重试。"""
|
||||
f = os.path.join(CACHE, f"{symbol}_{start}_{end}_{adjust}.csv")
|
||||
if os.path.exists(f):
|
||||
return pd.read_csv(f, parse_dates=["日期"])
|
||||
df = _fetch(symbol, start, end, adjust)
|
||||
df.to_csv(f, index=False, encoding="utf-8-sig")
|
||||
return df
|
||||
|
||||
def add_returns(df, price_col="收盘"):
|
||||
df = df.copy()
|
||||
df["ret"] = df[price_col].pct_change().fillna(0.0)
|
||||
return df
|
||||
|
||||
def net_value(ret):
|
||||
return (1 + pd.Series(ret).fillna(0.0)).cumprod()
|
||||
|
||||
def stats(ret, periods=252):
|
||||
ret = pd.Series(ret).fillna(0.0)
|
||||
nv = net_value(ret); total = nv.iloc[-1] - 1
|
||||
ann = (1 + total) ** (periods / max(len(ret), 1)) - 1
|
||||
vol = ret.std() * np.sqrt(periods)
|
||||
sharpe = (ret.mean() * periods) / vol if vol > 0 else np.nan
|
||||
dd = (nv / nv.cummax() - 1).min()
|
||||
calmar = ann / abs(dd) if dd < 0 else np.nan
|
||||
return {"总收益": f"{total:.2%}", "年化": f"{ann:.2%}", "年化波动": f"{vol:.2%}",
|
||||
"夏普": f"{sharpe:.2f}", "最大回撤": f"{dd:.2%}", "卡玛": f"{calmar:.2f}"}
|
||||
|
||||
def max_drawdown(nv):
|
||||
return (nv / nv.cummax() - 1).min()
|
||||
@@ -0,0 +1,10 @@
|
||||
#!/usr/bin/env bash
|
||||
# 一键跑练习:bash run.sh 或 bash run.sh 03
|
||||
cd "$(dirname "$0")"
|
||||
export PYTHONPATH="$(pwd)/../pylibs"
|
||||
export MPLCONFIGDIR="/tmp/mpl"
|
||||
mkdir -p ../data/cache
|
||||
case "$1" in
|
||||
01|02|03|04|05) python3 ${1}_*.py ;;
|
||||
*) for f in 0*.py; do echo; echo "############ $f ############"; python3 "$f"; done ;;
|
||||
esac
|
||||
@@ -0,0 +1,23 @@
|
||||
"""预热缓存:把练习需要的数据一次性拉齐(带重试+间隔)。之后离线可跑。"""
|
||||
import time, sys
|
||||
from lib_quant import get_daily, CACHE
|
||||
JOBS = [
|
||||
("600519","20200101","20241231"), ("600519","20230101","20241231"),
|
||||
("000001","20200101","20241231"), ("000001","20230101","20241231"),
|
||||
("300750","20200101","20241231"), ("300750","20230101","20241231"),
|
||||
("601318","20200101","20241231"), ("000858","20200101","20241231"),
|
||||
]
|
||||
for sym,s,e in JOBS:
|
||||
import os
|
||||
f=os.path.join(CACHE,f"{sym}_{s}_{e}_qfq.csv")
|
||||
if os.path.exists(f):
|
||||
print("cached", sym, s); continue
|
||||
for att in range(8):
|
||||
try:
|
||||
df=get_daily(sym,s,e); print("OK", sym, s, len(df)); break
|
||||
except Exception as ex:
|
||||
print("retry",sym,att,repr(ex)[:60]); time.sleep(5+att*3)
|
||||
else:
|
||||
print("FAIL", sym, s); sys.exit(1)
|
||||
time.sleep(3)
|
||||
print("ALL DONE")
|
||||
@@ -0,0 +1,38 @@
|
||||
# 视频候选清单(确认后再下载)
|
||||
|
||||
> 整理:小五 | 2026-09-28 | 来源:B站搜索实测(github/google 不通,走搜狗+必应国内)
|
||||
> 状态:**仅候选,未下载**。你勾选后我用 BBDown / yt-dlp 下到本地。
|
||||
|
||||
## 一、推荐下载(值回票价)
|
||||
|
||||
| # | 名称 | UP主 | 链接 | 类型 | 说明 |
|
||||
|---|---|---|---|---|---|
|
||||
| 1 | WorldQuant Brain《零基础学量化》AI版 第一课 | 水木人中 | https://www.bilibili.com/video/BV1d3oFBPES6 | 系统课 | 零基础成体系,从0讲策略 |
|
||||
| 2 | 《因子投资-方法与实践》石川 第1章 因子投资基础 | 量衍金工 | https://www.bilibili.com/video/BV1VxZcBqEsj | 读书课 | **石川《因子投资》**配套讲解,进阶重点 |
|
||||
| 3 | 《多因子策略实战》构建量化选股模型 | (投资经典精读) | https://www.bilibili.com/video/BV118at6ME3g | 读书课 | 多因子选股实战思路 |
|
||||
| 4 | 新手怎么开始学量化?认知篇:何为量化 | 龙猫说财 | https://www.bilibili.com/video/BV1Vtht62EPY | 入门 | 建立正确认知,避坑 |
|
||||
| 5 | 普通人如何从0开始做量化?(保姆级全流程) | QuantX西蒙斯 | https://www.bilibili.com/video/BV1JY8F61E6y | 入门 | 全流程走一遍 |
|
||||
| 6 | 5天学会 PTrade - AI协助策略 | 国金证券宜宾营业部 | https://www.bilibili.com/video/BV1aekKBkEVK | 券商课 | 券商官方,实盘环境入门 |
|
||||
| 7 | 怎样用 Deepseek Harness 做量化策略研究 | Mr看海 | https://www.bilibili.com/video/BV1Nzb367E5J | AI+量化 | 用LLM做研究,前沿 |
|
||||
| 8 | 【2026最新】Codex 量化实战教学 | 搞AI的阿黑 | https://www.bilibili.com/video/BV1jgtZ6fEHq | AI+量化 | 编程Agent做量化 |
|
||||
|
||||
## 二、可作补充(先看一/再决定)
|
||||
|
||||
| 名称 | UP主 | 链接 | 说明 |
|
||||
|---|---|---|---|
|
||||
| 横测:做量化最好用的大模型 | frank-quant | https://www.bilibili.com/video/BV1ijYB6QE1Y | 工具选型参考 |
|
||||
| 量化投资邢不行(系列) | 量化投资邢不行啊 | https://www.bilibili.com/video/BV1RZfxB1E68 | 拆解网红指标,防坑 |
|
||||
| 爆火Jev接入量化实盘实测 | 发明者量化 | https://www.bilibili.com/video/BV1kRez6BEuq | 实盘+AI |
|
||||
|
||||
## 三、非B站课程(文本/视频混合)
|
||||
- 慕课网 / 网易云课堂:搜「Python 量化投资」— 有系统付费课
|
||||
- 中国大学 MOOC:高校《金融工程》《量化投资》公开课(免费)
|
||||
|
||||
## 四、不推荐(已排除)
|
||||
- 币圈量化、跟单订阅、QMT连板超短、"永久躺平收益"类 —— 多为广告/割韭菜,已剔除
|
||||
|
||||
## 五、下载方案(确认后执行)
|
||||
- 工具:**BBDown**(B站专用,支持 1080P/4K、字幕、合集)+ **yt-dlp**(通用)
|
||||
- 输出:`量化学习/data/videos/`,可合并音频+视频、下字幕
|
||||
- ⚠️ 付费/会员课程需自备账号;我这边无账号只能下免费公开视频
|
||||
- 需要的话我可以顺带**提取字幕**,转成文本给你看(更符合你的文本偏好)
|
||||
@@ -0,0 +1,53 @@
|
||||
# 量化投资学习资料总表(适合国内环境)
|
||||
|
||||
> 汇总:小五 | 2026-09-28
|
||||
> 来源:小五主搜 + 小强(Gitea #31) + 小马(Gitea #30)(后者回帖后并入本表)
|
||||
> 说明:沙箱内 github/google 不通,链接优先给**国内可访问**的;GitHub 项目给 gitee 镜像路径。
|
||||
|
||||
## A. 书单(文本 · 打地基)
|
||||
|
||||
| 书名 | 作者 | 阶段 | 国内可购 | 备注 |
|
||||
|---|---|---|---|---|
|
||||
| 《打开量化投资的黑箱》第2版 | Rishi Narang | 入门 | ✅ 中文版 | 建立行业全貌,第一本 |
|
||||
| 《量化交易:如何建立自己的算法交易事业》 | Ernest Chan | 入门→实战 | ✅ 中文版 | 从0讲起,可照做 |
|
||||
| 《算法交易:制胜策略与原理》 | Ernest Chan | 进阶 | ✅ | 上本续作 |
|
||||
| 《Systematic Trading》 | Robert Carver | 进阶 | 英文为主 | 风险预算/系统化思想 |
|
||||
| 《主动投资组合管理》 | Grinold & Kahn | 进阶 | ✅ 中文版 | 因子/风险模型天花板 |
|
||||
| 《量化投资:策略与技术》 | 丁鹏 | 入门→进阶 | ✅ | 国内教材,A股语境 |
|
||||
| 《Python金融大数据分析》 | Yves Hilpisch | 工具 | ✅ | 配合实战 |
|
||||
| 《海龟交易法则》 | Curtis Faith | 入门 | ✅ | 趋势跟踪思想 |
|
||||
| 《统计套利》 | Andrew Pole | 进阶 | ✅ | 配对/均值回归 |
|
||||
|
||||
## B. 在线平台 / 开源框架(文本为主 · 国内可访问)
|
||||
|
||||
| 名称 | 类型 | 链接 | 阶段 | 备注 |
|
||||
|---|---|---|---|---|
|
||||
| 聚宽 JoinQuant | 在线研究+回测+社区 | https://www.joinquant.com | 全阶段 | **最省事入口**,A股数据全、策略帖海量 |
|
||||
| 米筐 RiceQuant | 在线平台 | https://www.ricequant.com | 全阶段 | 同类 |
|
||||
| 掘金量化 | 在线平台 | https://www.myquant.cn | 全阶段 | 同类 |
|
||||
| BigQuant | 在线AI量化 | https://bigquant.com | 进阶 | 机器学习向 |
|
||||
| AKShare | 免费数据源(库) | https://akshare.akfamily.xyz | 数据 | ✅ 中文文档,A股/期货/基金 |
|
||||
| Tushare | 数据源(库) | https://tushare.pro | 数据 | 需注册积分 |
|
||||
| backtrader | 回测框架 | 文档 https://www.backtrader.com (中文教程多) | 回测 | 轻量易上手 |
|
||||
| vnpy | 交易框架 | https://www.vnpy.com | 实盘 | 国内最主流,配书《vn.py 从入门到进阶》 |
|
||||
| Qlib(微软) | AI量化投研 | github.com/microsoft/qlib(用 gitee 镜像) | 进阶/ML | 因子+模型全流程 |
|
||||
|
||||
## C. 社区 / 专栏(文本)
|
||||
|
||||
| 名称 | 链接 | 备注 |
|
||||
|---|---|---|
|
||||
| 知乎「量化投资学习路线图·书籍篇」等专栏 | https://zhuanlan.zhihu.com | 路线梳理 |
|
||||
| CSDN 量化专栏 | https://blog.csdn.net | 中文实战多 |
|
||||
| 掘金 | https://juejin.cn | 工程向 |
|
||||
| 聚宽/米筐社区帖 | 平台内 | 策略源码 |
|
||||
|
||||
## D. 视频(候选见「视频候选清单.md」,确认后再下载)
|
||||
|
||||
## E. 按阶段的路线(建议照走)
|
||||
1. **基础**:Python + pandas/numpy + 金融常识 + 统计学
|
||||
2. **数据+回测**:AKShare/Tushare + 聚宽/backtrader → 写双均线、动量
|
||||
3. **策略进阶**:多因子、统计套利、CTA、事件驱动
|
||||
4. **工程/ML**:Qlib、因子库、ML选股、风控
|
||||
5. **实盘**:vnpy 接券商/期货
|
||||
|
||||
> 待办:小强(#31)、小马(#30) 回帖后并入本表并去重。
|
||||
+225
@@ -0,0 +1,225 @@
|
||||
# quanxiel 量化系统 · 完善路线图(ROADMAP)
|
||||
|
||||
> 编制:小五 | 2026-09-28
|
||||
> 定位:在现有「数据导入 + alpha 研究闭环」基础上,补齐 **回测可信度 / 风控 / 归因 / 自动化** 四块地基,打通「研究 → 实盘」。
|
||||
|
||||
---
|
||||
|
||||
## 0. 现状定位
|
||||
|
||||
| 已有 | 位置 |
|
||||
|---|---|
|
||||
| 数据导入(Tushare / 同花顺)+ 增量 + schema | `quantitative_data/` |
|
||||
| 数据加载(DB + 模拟回退) | `alpha/data_loader.py` |
|
||||
| 因子研究(注册表 / IC / 分层 / 中性化 / 换手) | `alpha/factors.py` |
|
||||
| 策略(信号生成 / 分位 / 组合打分 / 权重) | `alpha/strategy.py` |
|
||||
| 事件驱动回测(组合 / 持仓 / 交易成本) | `alpha/backtest.py` |
|
||||
| 绩效评估(收益 / 风险 / 报告) | `alpha/evaluation.py` |
|
||||
| 7 个策略 notebook | `alpha/*.ipynb` |
|
||||
|
||||
**结论**:研究闭环已成;「研究 → 实盘」的整条链路基本缺失。缺口见下表。
|
||||
|
||||
---
|
||||
|
||||
## 1. 模块矩阵(已有 / 不完整 / 缺失)
|
||||
|
||||
| 层 | 状态 | 说明 |
|
||||
|---|---|---|
|
||||
| 数据导入 | ✅ | 行情/财务/基本面/增量 |
|
||||
| 数据质检 | ⚠️ | 仅 `check_daily_coverage`,缺系统化 DQ |
|
||||
| Point-in-Time | ❌ | 财报未按公告日对齐;无成分股历史快照 |
|
||||
| 因子研究 | ✅ | 基础技术+基本面因子 |
|
||||
| 风格/风险模型 | ❌ | 无 Barra 类暴露与正交化 |
|
||||
| 组合构建 | ⚠️ | 有简单权重,无优化器/约束求解 |
|
||||
| 回测引擎 | ⚠️ | 无涨跌停/停牌/T+1/整手/参与率约束 |
|
||||
| 风控 | ❌ | 基本空白 |
|
||||
| 归因分析 | ❌ | 无 |
|
||||
| 稳健性验证 | ❌ | 无 WFO / PBO |
|
||||
| 实盘执行 | ❌ | 无交易接口/订单管理/对账 |
|
||||
| 工程运维 | ⚠️ | 有导入脚本,无调度/监控 |
|
||||
| 报告可视化 | ⚠️ | 有 markdown/df,无自动报告/dashboard |
|
||||
|
||||
---
|
||||
|
||||
## 2. P0 — 让研究「可信」(最先做)
|
||||
|
||||
> 不做这层,回测收益大概率是**假的**(未来函数 + 不可成交 + 无风控)。
|
||||
|
||||
### P0-1 回测真实性:A 股交易规则引擎
|
||||
- **新增** `alpha/market_rules.py`
|
||||
- **接口**:
|
||||
```python
|
||||
class MarketRules:
|
||||
"""A股交易规则约束,供回测下单前校验"""
|
||||
def can_buy(self, ts_code, date, price) -> bool: ... # 涨停/停牌/次新
|
||||
def can_sell(self, ts_code, date, price) -> bool: ... # 跌停/停牌
|
||||
def round_lot(self, qty, side) -> int: ... # 100股整手(卖出可零股)
|
||||
def max_volume(self, ts_code, date, participation=0.1) -> float: ... # 参与率上限
|
||||
def is_tradable(self, ts_code, date) -> bool: ... # 停牌过滤
|
||||
```
|
||||
- **集成**:`BacktestEngine(config, market_rules=MarketRules(...))`;下单前 `execute()` 走校验
|
||||
- **数据依赖**:`daily` 已有 `pre_close/pct_chg`,涨跌停按 ±10%/±20% 阈值 + `stock_basic.market` 判定;停牌用 `vol=0` 或 `trade_cal`
|
||||
- **验收**:构造一个"次日涨停买入"用例,断言该笔被拒
|
||||
|
||||
### P0-2 Point-in-Time(防未来函数)
|
||||
- **数据侧** `quantitative_data/schema.sql` 新增表 `index_weight`(成分股历史)
|
||||
```sql
|
||||
CREATE TABLE index_weight (
|
||||
index_code VARCHAR(20), con_code VARCHAR(20),
|
||||
trade_date DATE, weight NUMERIC(10,6),
|
||||
PRIMARY KEY (index_code, con_code, trade_date));
|
||||
```
|
||||
- **加载侧** `alpha/data_loader.py` 增加 `as_of` 语义
|
||||
```python
|
||||
def load_financials(self, fields, start, end, as_of=None):
|
||||
"""as_of 非空时,仅取 ann_date <= as_of 的财报(点对点快照)"""
|
||||
def load_universe(self, index_code, date) -> list[str]:
|
||||
"""取 date 当日在成分列表内的股票(历史成分)"""
|
||||
```
|
||||
- **验收**:同一财报在公告日前后取到不同值(公告前取不到)
|
||||
|
||||
### P0-3 风控限额(事前约束)
|
||||
- **新增** `alpha/risk/limits.py`
|
||||
```python
|
||||
class RiskLimits:
|
||||
def __init__(self, config: AlphaConfig): ...
|
||||
def clamp_weights(self, weights, industry_map) -> pd.Series: ...
|
||||
# 单票上限 max_position_pct / 行业上限 / 总仓位
|
||||
def check_drawdown_stop(self, nav_curve) -> bool: ... # 回撤熔断
|
||||
def risk_budget(self, weights, cov) -> dict: ... # 风险预算/贡献
|
||||
```
|
||||
- **集成**:回测每期调仓前 `clamp_weights`;净值触发回撤阈值→降仓/清仓
|
||||
- **验收**:给定超限组合,输出满足全部限额的权重
|
||||
|
||||
---
|
||||
|
||||
## 3. P1 — 提升质量
|
||||
|
||||
### P1-4 收益归因
|
||||
- **新增** `alpha/attribution.py`
|
||||
```python
|
||||
class BrinsonAttribution:
|
||||
def run(self, portfolio_weights, benchmark_weights,
|
||||
returns, industry_map) -> pd.DataFrame:
|
||||
"""返回 配置/选股/交互 三因子分解"""
|
||||
class FactorAttribution:
|
||||
def run(self, weights, factor_exposures, factor_returns) -> dict:
|
||||
"""因子收益贡献分解"""
|
||||
```
|
||||
|
||||
### P1-5 组合优化器
|
||||
- **新增** `alpha/optimizer.py`(依赖 `cvxpy`)
|
||||
```python
|
||||
class MeanVariance: def solve(self, mu, cov, constraints) -> weights
|
||||
class RiskParity: def solve(self, cov) -> weights
|
||||
class BlackLitterman: def solve(self, prior, views, ...) -> mu
|
||||
# constraints: 行业/风格中性、换手上限、个股上限、多空约束
|
||||
```
|
||||
- **验收**:约束可行域内求解,权重和=1,满足全部约束
|
||||
|
||||
### P1-6 稳健性 / 过拟合检验
|
||||
- **新增** `alpha/validation.py`
|
||||
```python
|
||||
class WalkForward: def run(self, strategy_factory, folds) -> pd.DataFrame # 滚动样本外
|
||||
def pbo(returns_matrix) -> float # 过拟合概率 (Bailey et al.)
|
||||
def deflated_sharpe(returns, n_trials) -> float
|
||||
def sensitivity(param_grid, evaluator) -> pd.DataFrame
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. P2 — 走向实盘
|
||||
|
||||
### P2-7 调度 + 监控
|
||||
- `ops/daily_pipeline.py`:盘后自动 增量导入 → 信号生成 → 风控 → 报告
|
||||
- 调度:群晖 DSM 计划任务 / cron;失败告警走 OpenClaw(webhook)
|
||||
- **新增** `ops/monitor.py`:数据到达检测、任务失败、仓位异常
|
||||
|
||||
### P2-8 实盘执行 + 对账
|
||||
- **新增** `execution/`
|
||||
```python
|
||||
class BrokerAdapter(ABC):
|
||||
def place_order(self, order) -> str: ...
|
||||
def cancel_order(self, order_id): ...
|
||||
def query_position(self) -> pd.DataFrame: ...
|
||||
class QmtAdapter(BrokerAdapter): ... # A股 QMT/miniQMT
|
||||
class SimAdapter(BrokerAdapter): ... # 仿真盘
|
||||
```
|
||||
- **对账**:每日 `query_position` vs 内部账本,差异告警
|
||||
|
||||
### P2-9 报告 + Dashboard
|
||||
- **新增** `report/`:自动日报/周报(净值/回撤/持仓/归因/暴露)
|
||||
- Dashboard:复用 gemdalepi 静态站点方式,生成净值/回撤/暴露可视化页
|
||||
|
||||
---
|
||||
|
||||
## 5. 目录规划(目标)
|
||||
|
||||
```
|
||||
quanxiel/
|
||||
├── quantitative_data/ # 数据(现有)+ index_weight
|
||||
├── alpha/ # 研究(现有)
|
||||
│ ├── market_rules.py # [P0-1] 新增
|
||||
│ ├── risk/ # [P0-3] 新增
|
||||
│ ├── attribution.py # [P1-4] 新增
|
||||
│ ├── optimizer.py # [P1-5] 新增
|
||||
│ └── validation.py # [P1-6] 新增
|
||||
├── execution/ # [P2-8] 新增
|
||||
├── report/ # [P2-9] 新增
|
||||
└── ops/ # 运维(现有)+ daily_pipeline.py / monitor.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. 建议实施顺序
|
||||
|
||||
1. **P0-1 交易规则** → 2. **P0-2 PIT** → 3. **P0-3 风控限额**
|
||||
4. P1-4 归因 → 5. P1-5 优化器 → 6. P1-6 稳健性
|
||||
7. P2-7 调度 → 8. P2-8 实盘 → 9. P2-9 报告
|
||||
|
||||
> 交付节奏建议:每个模块「接口 + 最小实现 + 单测/验收用例」三件套,合入前跑通验收。
|
||||
|
||||
---
|
||||
|
||||
## 7. 版本 v1.1 — 书库佐证与增补(2026-09-28)
|
||||
|
||||
> 依据自家 talebook 书库 8 本量化书交叉验证本路线图,补充以下模块与验收项。
|
||||
> 参考书(talebook id):打开量化投资的黑箱(155) · Quantitative Trading, Chan(986) · Python for Algorithmic Trading, Hilpisch(1125) · 主动投资组合管理, Grinold&Kahn(183) · 量化交易之路, 阿布(149) · Python量化交易教程(167) · Quantitative Trading: Algorithms... Guo(983) · 量化投资策略:超额收益Alpha(1202)
|
||||
|
||||
### 7.1 目标架构对齐「黑箱」六模型
|
||||
《黑箱》(155) 给出交易系统标准结构:**Alpha 模型 + 风险模型 + 交易成本模型 → 投资组合构建模型 ↔ 执行模型**(另含 数据 / 研究)。
|
||||
|
||||
| 黑箱模块 | 本项目对应 | 状态 |
|
||||
|---|---|---|
|
||||
| Alpha 模型 | `alpha/factors.py` + `alpha/strategy.py` | ✅ |
|
||||
| 风险模型 | 需新增 `alpha/risk/risk_model.py`(因子协方差/暴露) | ❌ |
|
||||
| 交易成本模型 | `alpha/config.py` 有费率,缺**滑点/市场冲击**模型 | ⚠️ |
|
||||
| 组合构建模型 | 有简单权重,缺**优化器** | ⚠️ |
|
||||
| 执行模型 | 需新增 `execution/` | ❌ |
|
||||
|
||||
> 结论:现有目录只覆盖了黑箱 6 模型中的 **Alpha +(半)组合构建**;风险模型、成本模型、执行模型需显式补齐。
|
||||
|
||||
### 7.2 新增模块(在原 P0/P1/P2 上增补)
|
||||
- **P0-4 数据质检细化**(986:数据是否复权/是否 survivorship-bias free)→ 新增 `quantitative_data/dq.py`:①复权口径核对(除权除息) ②幸存者偏差检查(point-in-time 股票池)③高低价数据口径 ④停牌/退市标记
|
||||
- **P1-7 资金管理 / 仓位分配**(149 凯利公式;986 Optimal Capital Allocation)→ 新增 `alpha/money_mgmt.py`:Kelly / 固定比例 / 波动率目标;对应"避免重仓"
|
||||
- **P1-8 因子处理流水线显式化**(167 优矿 RDP:去极值/中性化/标准化)→ 新增 `alpha/preprocess.py`:`winsorize()` / `zscore()` / `neutralize(industry, size)`,做成可复用管线(`factors.py` 已有部分,建议独立)
|
||||
- **P2-10 部署工程化**(1125 副标题即 *From Idea to Cloud Deployment*)→ Docker 化 + 依赖锁定 + CI(单测) + 定时任务容器化
|
||||
- **P2-11 纸面交易 / 仿真盘**(986:Paper Trading 是"最终的样本外检验")→ `execution/SimAdapter` + 与回测结果对比报告,**建议提前为 P2 第一项**
|
||||
|
||||
### 7.3 方法论 / 验收项增补
|
||||
- **回测偏差自检清单**(986 第3章 / 1125 第4章):前视偏差 · 数据窥探(data-snooping) · 交易成本 · 幸存者偏差 · 高低价数据 —— 作为回测模块验收 checklist
|
||||
- **样本外 + 参数近邻稳定性**(155/149):网格寻优后检查近邻参数结果是否相近;不相近 ⇒ 疑似过拟合
|
||||
- **数据挖掘四方针**(183):直觉 · 克制 · 合乎情理 · 样本外测试
|
||||
- **IC→Alpha 基本定律**(183 预测基本定理):`α = IC × 波动率 × 标准化信号`,`IR = IC × √广度`;建议在 P1-5 优化器上游加入此转换,用于衡量研究"深度 vs 广度"
|
||||
|
||||
### 7.4 修订后的实施顺序
|
||||
1. **P0-4 数据质检** → 2. P0-1 交易规则 → 3. P0-2 PIT → 4. P0-3 风控
|
||||
5. **P1-8 因子管线** → 6. P1-4 归因 → 7. **P1-7 资金管理** → 8. P1-5 优化器 → 9. P1-6 稳健性
|
||||
10. **P2-11 仿真盘** → 11. **P2-10 部署工程化** → 12. P2-7 调度监控 → 13. P2-8 实盘 → 14. P2-9 报告
|
||||
|
||||
### 7.5 主要「新发现」的缺口(原 ROADMAP 未列)
|
||||
1. **资金管理/仓位分配**(凯利/波动率目标)—— 原稿完全没有
|
||||
2. **交易成本模型细化**到滑点 + 市场冲击(不只费率)
|
||||
3. **显式风险模型**(因子协方差/暴露),而非仅在 config 里放限额
|
||||
4. **纸面交易**作为回测→实盘之间的强制关卡
|
||||
5. **部署工程化**(Docker/CI/云)—— 原稿只在 P2 一句带过
|
||||
@@ -0,0 +1,12 @@
|
||||
# 知识库 · Anki 牌组
|
||||
|
||||
`量化基础.apkg` —— 与 wiki「🧠 知识」板块一一对应的名词卡(36 张)。
|
||||
|
||||
- 导入方法见 wiki:[[知识/复习与Anki说明]]
|
||||
- **Android**:AnkiDroid → ⋮ → 导入 → 选此文件
|
||||
- 牌组名:`量化基础 · quanxiel`
|
||||
- 每张卡两种题型:①名词→详解 ②定义→名词
|
||||
- 复习节奏:1/2/4/7/15 天(间隔重复由 Anki 自动安排)
|
||||
|
||||
> 重新生成:更新词条后由小五用 genanki 重新打包(脚本见对话记录)。
|
||||
> 新增名词:直接告诉小五「加一个 XXX 名词卡」。
|
||||
Binary file not shown.
@@ -0,0 +1,30 @@
|
||||
# ops/ — 群晖主机侧运维脚本
|
||||
|
||||
## openclaw-webhook-iptables.sh
|
||||
|
||||
让 **Gitea 的 webhook 能送达 OpenClaw 容器**(`172.21.0.2:8899`)。
|
||||
|
||||
- **作用**:在群晖主机上维护两条 iptables 规则
|
||||
1. `nat PREROUTING`:主机 8899 → 容器 172.21.0.2:8899(DNAT)
|
||||
2. `DOCKER-USER`:放行跨 docker 网络转发(插到最前,绕过隔离 DROP)
|
||||
- **幂等**:重复运行不会产生重复规则
|
||||
- **需 root**;开机时自动等待 `DOCKER-USER` 链就绪(最多 ~3 分钟)
|
||||
|
||||
### 部署
|
||||
```sh
|
||||
sudo mkdir -p /volume1/scripts
|
||||
sudo cp ops/openclaw-webhook-iptables.sh /volume1/scripts/
|
||||
sudo chmod +x /volume1/scripts/openclaw-webhook-iptables.sh
|
||||
```
|
||||
DSM → 控制面板 → 任务计划 → 新增(触发的任务)→ 用户账号 root → 计划「开机启动」→ 运行命令:
|
||||
```
|
||||
sh /volume1/scripts/openclaw-webhook-iptables.sh
|
||||
```
|
||||
|
||||
### 变量
|
||||
- `OC_IP`:OpenClaw 容器 IP(默认 `172.21.0.2`)。容器重建后 IP 若变化需同步修改。
|
||||
- `PORT`:接收服务端口(默认 `8899`)
|
||||
|
||||
### 相关
|
||||
- 排障全过程与最终方案见 issue **#24**
|
||||
- webhook 目标 URL:`http://192.168.27.11:8899/gitea`
|
||||
@@ -0,0 +1,55 @@
|
||||
# course-toolkit — YouTube 课程「下载 + 转写 + 成册」工具包
|
||||
|
||||
把某个 YouTube 课程/播放列表,批量变成 **本地视频 + 英文文稿 + 中文导读 Markdown**。
|
||||
|
||||
## 组成
|
||||
| 文件 | 作用 |
|
||||
|---|---|
|
||||
| `fetch_course.sh` | 主入口:枚举播放列表 → 并行下载+转写 → 清洗文稿 |
|
||||
| `worker.sh` | 单个视频的下载/抽音频/转写(被主脚本并行调用) |
|
||||
| `course_util.py` | 子命令:`manifest` / `clean` / `transcribe` / `build` |
|
||||
|
||||
## 快速开始
|
||||
```bash
|
||||
# 1) 下载 + 转写(默认英文,4 路并行)
|
||||
./fetch_course.sh "https://www.youtube.com/playlist?list=XXXX" ~/course/我的课程
|
||||
# 中文课程: MODEL=small LANG=zh ./fetch_course.sh <url> <outdir>
|
||||
|
||||
# 2) 生成中文导读:在 <outdir>/guides/ 下放 01.txt 02.txt ...(可直接让 AI 读 clean/NN.txt 生成)
|
||||
# 然后组装成册:
|
||||
python3 course_util.py build ~/course/我的课程
|
||||
```
|
||||
|
||||
产物结构:
|
||||
```
|
||||
<outdir>/
|
||||
├── manifest.txt # NN|id|slug|title
|
||||
├── videos/NN-slug.mp4 # 本地视频(默认 360p)
|
||||
├── transcripts/ # 带时间轴文稿
|
||||
├── clean/NN.txt # 去时间轴文稿(供检索/摘要)
|
||||
├── guides/NN.txt # 中文导读(可选,人工/AI 写)
|
||||
├── NN-slug.md # 成册笔记(导读 + 全文)
|
||||
├── README.md # 目录索引
|
||||
└── build.log # 运行日志
|
||||
```
|
||||
|
||||
## 环境变量
|
||||
| 变量 | 默认 | 说明 |
|
||||
|---|---|---|
|
||||
| `PROXY` | `socks5://192.168.27.96:1080` | 访问 YouTube 的代理 |
|
||||
| `MODEL` | `base.en` | faster-whisper 模型(英文 base.en/small.en;中文 small)|
|
||||
| `LANG` | `en` | 转写语言 |
|
||||
| `JOBS` | `4` | 并行数 |
|
||||
| `FMT` | `18` | yt-dlp 格式(见下)|
|
||||
|
||||
依赖:`yt-dlp`、`ffmpeg`、`faster-whisper` + `python3`。
|
||||
|
||||
## 关键经验(踩过的坑)
|
||||
1. **画质问题**:YouTube 未登录下载被限,实测只有 `--extractor-args "youtube:player_client=android"` 能下 → **上限 360p**(格式 18,渐进式、含音轨)。`android_vr`/`tv_embedded` 会列出 1080p 但下载 **403**;要高清需 PO token / 登录 cookie。
|
||||
2. **官方字幕不稳**:常遇 **429 限流**或该视频**根本没有字幕** → 故用**本地 ASR**替代。
|
||||
3. **模型下载**:走 `hf-mirror`,并必须 `HF_HUB_DISABLE_XET=1`(否则 xet CAS 401)。
|
||||
4. **速度**:faster-whisper `base.en` + `beam_size=1` + **单线程多进程并行**,实测 ~28x 实时(4 并行);19 集≈12 分钟。
|
||||
5. **成册**:每集 md = 中文导读 + 英文全文;文稿为自动转写,**未人工校对**。
|
||||
6. ⚠️ 版权:视频版权归原作者,建议仅作**个人学习**归档,勿公开分发。
|
||||
|
||||
_(本工具包由小五沉淀,首个用户:《Full Algorithmic Trading Using Python》19 集)_
|
||||
@@ -0,0 +1,102 @@
|
||||
#!/usr/bin/env python3
|
||||
"""course-toolkit 辅助工具。
|
||||
|
||||
子命令:
|
||||
manifest <rawfile> 将 index|id|title 转为 NN|id|slug|title
|
||||
clean <outdir> transcripts/*.txt -> clean/*.txt(去时间轴、并段)
|
||||
transcribe <wav> <out> 本地 ASR 转写(env: MODEL, LANG, THREADS)
|
||||
build <outdir> 组装 md + README(guides/<NN>.txt 可选作中文导读)
|
||||
"""
|
||||
import os, re, sys
|
||||
|
||||
def slugify(s):
|
||||
s = s.lower()
|
||||
s = re.sub(r'[^a-z0-9]+', '-', s)
|
||||
return re.sub(r'-+', '-', s).strip('-')[:60] or 'video'
|
||||
|
||||
def cmd_manifest(rawfile):
|
||||
out = []
|
||||
for line in open(rawfile, encoding='utf-8'):
|
||||
line = line.strip()
|
||||
if not line: continue
|
||||
parts = line.split('|')
|
||||
if len(parts) < 3: continue
|
||||
idx, vid, title = parts[0], parts[1], '|'.join(parts[2:])
|
||||
try: nn = f"{int(idx):02d}"
|
||||
except Exception: nn = slugify(idx)[:2]
|
||||
out.append(f"{nn}|{vid}|{slugify(title)}|{title}")
|
||||
sys.stdout.write("\n".join(out) + "\n")
|
||||
|
||||
def cmd_clean(outdir):
|
||||
tdir = os.path.join(outdir, 'transcripts'); cdir = os.path.join(outdir, 'clean')
|
||||
os.makedirs(cdir, exist_ok=True)
|
||||
for f in sorted(os.listdir(tdir)):
|
||||
if not f.endswith('.txt'): continue
|
||||
nn = f.split('-')[0]
|
||||
lines = [l for l in open(os.path.join(tdir, f), encoding='utf-8').read().splitlines() if l.strip()]
|
||||
txt = " ".join(re.sub(r'^\[\s*[\d.]+\s*-\s*[\d.]+\s*\]\s*', '', l) for l in lines)
|
||||
open(os.path.join(cdir, f"{nn}.txt"), 'w', encoding='utf-8').write(txt)
|
||||
print(f"[clean] {len(os.listdir(cdir))} files -> {cdir}")
|
||||
|
||||
def cmd_transcribe(wav, out):
|
||||
os.environ.setdefault("HF_ENDPOINT", "https://hf-mirror.com")
|
||||
os.environ.setdefault("HF_HUB_DISABLE_XET", "1")
|
||||
from faster_whisper import WhisperModel
|
||||
model = os.environ.get("MODEL", "base.en")
|
||||
lang = os.environ.get("LANG", "en")
|
||||
threads = int(os.environ.get("THREADS", "1"))
|
||||
m = WhisperModel(model, device="cpu", compute_type="int8", cpu_threads=threads)
|
||||
segs, info = m.transcribe(wav, language=lang, beam_size=1, vad_filter=True)
|
||||
lines = [f"[{s.start:7.1f}-{s.end:7.1f}] {s.text.strip()}" for s in segs]
|
||||
open(out, 'w', encoding='utf-8').write("\n".join(lines))
|
||||
print(f"[transcribe] {out} segs={len(lines)} dur={info.duration:.0f}")
|
||||
|
||||
def cmd_build(outdir):
|
||||
import subprocess, json
|
||||
man = [l.split('|') for l in open(os.path.join(outdir, 'manifest.txt'), encoding='utf-8').read().splitlines() if l.strip()]
|
||||
gdir = os.path.join(outdir, 'guides')
|
||||
rows = []
|
||||
for p in man:
|
||||
NN, ID, SLUG = p[0], p[1], p[2]
|
||||
title = p[3] if len(p) > 3 else SLUG
|
||||
vid = f"videos/{NN}-{SLUG}.mp4"
|
||||
txt = open(os.path.join(outdir, 'clean', f'{NN}.txt'), encoding='utf-8').read().strip()
|
||||
gpath = os.path.join(gdir, f'{NN}.txt')
|
||||
guide = open(gpath, encoding='utf-8').read().strip() if os.path.exists(gpath) else "> TODO:待补充中文导读。"
|
||||
try:
|
||||
dur = float(subprocess.check_output(["ffprobe", "-v", "error", "-show_entries", "format=duration", "-of", "csv=p=0", os.path.join(outdir, vid)]).decode())
|
||||
mm, ss = int(dur // 60), int(dur % 60)
|
||||
except Exception:
|
||||
mm = ss = 0
|
||||
md = f"""# {NN} · {title}
|
||||
|
||||
- **原始视频**:https://youtu.be/{ID}
|
||||
- **时长**:{mm} 分 {ss} 秒
|
||||
- **本地视频**:[{vid}]({vid})
|
||||
|
||||
## 🎯 本集要点(中文导读)
|
||||
|
||||
{guide}
|
||||
|
||||
## 📝 完整文稿(自动转写)
|
||||
|
||||
> 由本地 ASR 转写,未人工校对,供检索/精读使用。
|
||||
|
||||
{txt}
|
||||
"""
|
||||
open(os.path.join(outdir, f'{NN}-{SLUG}.md'), 'w', encoding='utf-8').write(md)
|
||||
rows.append((NN, title, mm, ss, SLUG))
|
||||
idx = "\n".join(f"| {n} | {t} | {m}:{s:02d} | [📄 笔记]({n}-{sl}.md) · [🎬 视频](videos/{n}-{sl}.mp4) |" for n, t, m, s, sl in rows)
|
||||
open(os.path.join(outdir, 'README.md'), 'w', encoding='utf-8').write(
|
||||
"# 课程教程\n\n> 由 course-toolkit 生成。\n\n| 集 | 标题 | 时长 | 链接 |\n|---|---|---|---|\n" + idx + "\n")
|
||||
print(f"[build] {len(rows)} 篇 + README -> {outdir}")
|
||||
|
||||
if __name__ == '__main__':
|
||||
if len(sys.argv) < 2:
|
||||
print(__doc__); sys.exit(1)
|
||||
cmd = sys.argv[1]
|
||||
if cmd == 'manifest': cmd_manifest(sys.argv[2])
|
||||
elif cmd == 'clean': cmd_clean(sys.argv[2])
|
||||
elif cmd == 'transcribe': cmd_transcribe(sys.argv[2], sys.argv[3])
|
||||
elif cmd == 'build': cmd_build(sys.argv[2])
|
||||
else: print(__doc__); sys.exit(1)
|
||||
@@ -0,0 +1,34 @@
|
||||
#!/bin/bash
|
||||
# fetch_course.sh — 下载 YouTube 课程视频 + 本地转写 + 组装教程
|
||||
#
|
||||
# 用法:
|
||||
# ./fetch_course.sh <playlist_url或单个视频url> <输出目录>
|
||||
#
|
||||
# 环境变量(可选):
|
||||
# PROXY 代理,默认 socks5://192.168.27.96:1080
|
||||
# MODEL whisper 模型,默认 base.en(英文用 base.en/small.en,中文用 small)
|
||||
# LANG 语言,默认 en(中文用 zh)
|
||||
# JOBS 并行数,默认 4
|
||||
# FMT yt-dlp 格式,默认 18(360p 渐进式,含音轨;见 README 画质说明)
|
||||
set -uo pipefail
|
||||
URL="${1:?用法: ./fetch_course.sh <url> <输出目录>}"
|
||||
OUT="${2:?缺少输出目录}"
|
||||
PROXY="${PROXY:-socks5://192.168.27.96:1080}"
|
||||
MODEL="${MODEL:-base.en}"; LG="${LANG:-en}"; JOBS="${JOBS:-4}"; FMT="${FMT:-18}"
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
mkdir -p "$OUT/videos" "$OUT/transcripts" "$OUT/clean"
|
||||
|
||||
if [ ! -s "$OUT/manifest.txt" ]; then
|
||||
echo "[*] 枚举视频列表..."
|
||||
yt-dlp --proxy "$PROXY" --no-warnings --flat-playlist \
|
||||
--print "%(playlist_index)s|%(id)s|%(title)s" "$URL" > "$OUT/_raw.txt" || { echo "枚举失败"; exit 1; }
|
||||
python3 "$HERE/course_util.py" manifest "$OUT/_raw.txt" > "$OUT/manifest.txt"
|
||||
fi
|
||||
echo "[*] 待处理视频: $(grep -c . "$OUT/manifest.txt") 个 | 模型=$MODEL 并行=$JOBS"
|
||||
|
||||
export PROXY MODEL LANG="$LG" OUT HERE FMT
|
||||
grep -v '^[[:space:]]*$' "$OUT/manifest.txt" | xargs -P "$JOBS" -I{} "$HERE/worker.sh" "{}"
|
||||
|
||||
python3 "$HERE/course_util.py" clean "$OUT"
|
||||
echo "[*] 全部完成: $OUT"
|
||||
echo "[*] 下一步: 在 $OUT/guides/ 放中文导读(<NN>.txt)后运行 course_util.py build $OUT"
|
||||
@@ -0,0 +1,23 @@
|
||||
#!/bin/bash
|
||||
# worker.sh — 处理单个视频:下载 → 抽音频 → 转写(由 fetch_course.sh 调用)
|
||||
set -uo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
line="$1"
|
||||
NN=$(echo "$line"|cut -d'|' -f1); ID=$(echo "$line"|cut -d'|' -f2); SLUG=$(echo "$line"|cut -d'|' -f3)
|
||||
VID="$OUT/videos/${NN}-${SLUG}.mp4"; TXT="$OUT/transcripts/${NN}-${SLUG}.${LANG}.txt"
|
||||
LOG="$OUT/build.log"
|
||||
log(){ echo "[$(date '+%F %T')] [$NN] $*" >> "$LOG"; }
|
||||
[ -f "$TXT" ] && { log "skip(done)"; exit 0; }
|
||||
if [ ! -f "$VID" ]; then
|
||||
for a in 1 2 3; do
|
||||
yt-dlp --proxy "$PROXY" --no-warnings --no-part \
|
||||
--extractor-args "youtube:player_client=android" -f "$FMT" \
|
||||
-o "$VID" "https://www.youtube.com/watch?v=$ID" >>"$LOG" 2>&1 && break
|
||||
log "retry-dl $a"; sleep 8
|
||||
done
|
||||
fi
|
||||
[ -f "$VID" ] || { log "FAIL download"; exit 1; }
|
||||
ffmpeg -y -hide_banner -loglevel error -i "$VID" -ar 16000 -ac 1 -c:a pcm_s16le "$OUT/_${NN}.wav" 2>>"$LOG" || { log "FAIL audio"; exit 1; }
|
||||
MODEL="$MODEL" LANG="$LANG" THREADS=1 python3 "$HERE/course_util.py" transcribe "$OUT/_${NN}.wav" "$TXT" >>"$LOG" 2>&1 || { log "FAIL transcribe"; exit 1; }
|
||||
rm -f "$OUT/_${NN}.wav"
|
||||
log "ok $(du -h "$VID"|cut -f1)"
|
||||
@@ -0,0 +1,47 @@
|
||||
#!/bin/sh
|
||||
# ============================================================
|
||||
# OpenClaw Gitea webhook 通道:群晖主机 8899 -> OpenClaw 容器
|
||||
# 幂等:重复运行不会产生重复规则;需 root 运行
|
||||
# 适用:群晖 DSM 7.x,任务计划-开机启动
|
||||
# ============================================================
|
||||
export PATH=/sbin:/bin:/usr/sbin:/usr/bin:$PATH
|
||||
|
||||
OC_IP="172.21.0.2" # OpenClaw 容器 IP(容器重建后若 IP 变了,改这里)
|
||||
PORT="8899" # 接收服务端口
|
||||
LOG="/volume1/scripts/openclaw-webhook-iptables.log"
|
||||
|
||||
mkdir -p "$(dirname "$LOG")" 2>/dev/null
|
||||
log() { echo "[$(date '+%F %T')] $*" >> "$LOG"; }
|
||||
|
||||
log "=== run start (OC_IP=$OC_IP PORT=$PORT) ==="
|
||||
|
||||
# --- 1) DNAT:群晖主机 PORT -> OpenClaw 容器 ---
|
||||
if iptables -t nat -C PREROUTING -p tcp --dport "$PORT" \
|
||||
-j DNAT --to-destination "$OC_IP:$PORT" 2>/dev/null; then
|
||||
log "DNAT 已存在,跳过"
|
||||
else
|
||||
iptables -t nat -A PREROUTING -p tcp --dport "$PORT" \
|
||||
-j DNAT --to-destination "$OC_IP:$PORT" && log "已添加 DNAT"
|
||||
fi
|
||||
|
||||
# --- 2) 放行跨 docker 网络转发(插到 DOCKER-USER 最前,绕过隔离 DROP)---
|
||||
# 开机时 docker 可能还没起,先等 DOCKER-USER 链就绪
|
||||
i=0
|
||||
while ! iptables -L DOCKER-USER -n >/dev/null 2>&1; do
|
||||
i=$((i+1))
|
||||
[ "$i" -gt 36 ] && { log "等待 DOCKER-USER 超时,放弃"; exit 1; }
|
||||
sleep 5
|
||||
done
|
||||
|
||||
if iptables -C DOCKER-USER -p tcp -d "$OC_IP" --dport "$PORT" -j ACCEPT 2>/dev/null; then
|
||||
log "DOCKER-USER 放行已存在,跳过"
|
||||
else
|
||||
iptables -I DOCKER-USER 1 -p tcp -d "$OC_IP" --dport "$PORT" -j ACCEPT \
|
||||
&& log "已添加 DOCKER-USER 放行"
|
||||
fi
|
||||
|
||||
log "=== run done ==="
|
||||
|
||||
# 可选的即时自检(注释掉即可)
|
||||
# iptables -t nat -L PREROUTING -n | grep ":$PORT"
|
||||
# iptables -L DOCKER-USER -n | grep "$OC_IP"
|
||||
@@ -0,0 +1,168 @@
|
||||
# 同花顺金融数据服务(hithink-finance)接入指南
|
||||
|
||||
> 同花顺官方 A 股金融数据服务,一个 API Key 打通行情/财报/估值/特色数据。
|
||||
> 官网:https://fuyao.aicubes.cn | GitHub:https://github.com/HiThink-Tech/Financial-API(开源工具链)
|
||||
> 更新:2026-08-22 | 状态:✅ 已实测可用(沙箱已配置)
|
||||
|
||||
## 1. 概述
|
||||
|
||||
面向 AI Agent、量化研究和应用开发的 A 股数据服务,覆盖:
|
||||
|
||||
- **行情**:实时快照、历史 K 线、复权因子、公司行动、交易日历
|
||||
- **财务**:利润表、资产负债表、现金流量表、财务指标
|
||||
- **估值**:市盈率 TTM/MRQ、市净率、市销率、市现率
|
||||
- **特色数据**:集合竞价、涨跌停池、炸板池、连板天梯、个股异动、热榜、龙虎榜
|
||||
- **指数/板块**:目录、成分股、行情、历史 K 线
|
||||
- **公募基金**:资料、公司、经理、财务、持仓、净值、业绩、场内行情
|
||||
- **全市场导出**:全量/增量日 K、公司行动等标准数据文件
|
||||
|
||||
**明确不覆盖**:分钟 K、tick、海外行情、宏观数据、新闻公告原文、研报。
|
||||
(请求未支持的数据时明确说明,不得用模拟数据冒充。)
|
||||
|
||||
## 2. API Key
|
||||
|
||||
- 申请:https://fuyao.aicubes.cn/admin/ 一键签发
|
||||
- 统一环境变量:`HITHINK_FINANCE_API_KEY`(API / MCP / CLI / Python 共用一把 Key)
|
||||
- **安全要求**:Key 只写入用户级凭据来源(credentials.env / CLI 凭据库),
|
||||
**严禁**写入代码、日志、公开配置或 Git 仓库;Agent 不得复述 Key。
|
||||
|
||||
家庭环境现状(2026-08-22 已配好):
|
||||
- 沙箱(小五所在环境):`~/.openclaw/credentials.env`(600 权限)+ CLI 凭据库
|
||||
- 小强如需独立使用:向爸爸要 Key,按 4.1 节 `auth login` 配置自己的凭据
|
||||
|
||||
## 3. 接入方式速览
|
||||
|
||||
| 场景 | 推荐方式 |
|
||||
|------|---------|
|
||||
| Agent 自动查数 | hithink-finance Skill(npx skills add HiThink-Tech/Financial-API --skill hithink-finance -g --yes) |
|
||||
| Claude/Cursor 对话查数 | MCP(4 个托管端点,见 4.4) |
|
||||
| Python/Notebook 研究 | Python SDK(见 4.3) |
|
||||
| 网站/App/后端接入 | REST API(见 4.5) |
|
||||
| 终端批量查询/导出 | CLI(见 4.1) |
|
||||
| 本地长期保存 + SQL 研究 | marketdb / 本地 DuckDB(见 4.2) |
|
||||
|
||||
## 4. 详细接入
|
||||
|
||||
### 4.1 CLI(沙箱已装 ✅)
|
||||
|
||||
```bash
|
||||
npm install -g @hithink-tech/hithink-finance-cli --registry=https://registry.npmmirror.com
|
||||
hithink-finance auth login # 录入 API Key(--api-key-stdin 支持 stdin)
|
||||
hithink-finance capabilities --format json # 查看本版本能力目录
|
||||
```
|
||||
|
||||
常用命令(全部支持 `--format json` 稳定输出):
|
||||
|
||||
```bash
|
||||
# 按代码/名称/关键词找标的(返回唯一 thscode)
|
||||
hithink-finance symbol search --q 600519 --limit 5 --format json
|
||||
|
||||
# 最新行情快照(单只/多只/全市场)
|
||||
hithink-finance market snapshot --thscodes 600519.SH --format json
|
||||
|
||||
# 历史 K 线
|
||||
hithink-finance market history --thscode 600519.SH --kline daily --limit 30 --format json
|
||||
|
||||
# 财务报表(最近 N 期)
|
||||
hithink-finance financials income --thscode 600519.SH --limit 4 --format json
|
||||
|
||||
# 初始化本地 DuckDB(回测数据库)
|
||||
hithink-finance data init --format json
|
||||
|
||||
# SQL 查询本地前复权日线
|
||||
hithink-finance db query --sql "SELECT * FROM v_daily_qfq LIMIT 10" --format json
|
||||
```
|
||||
|
||||
> 具体命令/参数以 `hithink-finance capabilities --format json` 返回的机器可读能力目录为准。
|
||||
|
||||
### 4.2 本地 DuckDB(marketdb)—— 回测数据方案
|
||||
|
||||
```bash
|
||||
hithink-finance data init # 初始化本地库
|
||||
hithink-finance data sync # 增量同步(可定时)
|
||||
hithink-finance db query --sql "SELECT ..." --format json # 只读 SQL
|
||||
# 导出:data export / dump 相关命令(见 capabilities)
|
||||
```
|
||||
|
||||
适合:长期保存历史行情、复权计算、SQL 因子研究、回测数据集。
|
||||
替代/互补现有 quantitative_data 的 Tushare 导入链路(importer.py / incremental_import.py)。
|
||||
|
||||
### 4.3 Python SDK
|
||||
|
||||
```bash
|
||||
git clone https://github.com/HiThink-Tech/Financial-API
|
||||
cd Financial-API && pip install -e ./python
|
||||
```
|
||||
|
||||
```bash
|
||||
python python/toolkit/fuyao/scripts/fuyao.py tickers-search --q "贵州茅台"
|
||||
python python/toolkit/fuyao/scripts/fuyao.py prices-snapshot --thscodes 600519.SH
|
||||
```
|
||||
|
||||
### 4.4 MCP(Claude Desktop / Cursor / Windsurf)
|
||||
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"hithink-finance-a-share": {
|
||||
"type": "http",
|
||||
"url": "https://fuyao.aicubes.cn/mcp/a-share",
|
||||
"headers": { "X-api-key": "${HITHINK_FINANCE_API_KEY}" }
|
||||
},
|
||||
"hithink-finance-a-share-index": {
|
||||
"type": "http",
|
||||
"url": "https://fuyao.aicubes.cn/mcp/a-share-index",
|
||||
"headers": { "X-api-key": "${HITHINK_FINANCE_API_KEY}" }
|
||||
},
|
||||
"hithink-finance-meta": {
|
||||
"type": "http",
|
||||
"url": "https://fuyao.aicubes.cn/mcp/meta",
|
||||
"headers": { "X-api-key": "${HITHINK_FINANCE_API_KEY}" }
|
||||
},
|
||||
"hithink-finance-fund": {
|
||||
"type": "http",
|
||||
"url": "https://fuyao.aicubes.cn/mcp/fund",
|
||||
"headers": { "X-api-key": "${HITHINK_FINANCE_API_KEY}" }
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 4.5 REST API
|
||||
|
||||
- Base:`https://fuyao.aicubes.cn`,鉴权 Header:`X-api-key`
|
||||
- 统一 ApiResponse 信封:业务结果(含错误)HTTP 200 返回,用 `code` 字段分发
|
||||
- 路径风格:`/api/<标的宇宙>/<数据类型>/<动作>`,如 `/api/a-share/prices/snapshot`
|
||||
- snake_case 字段、显式 currency、毫秒级 Unix 时间戳(LLM-friendly)
|
||||
- 完整契约:https://fuyao.aicubes.cn/llms-full.txt | 仓库 docs/api/
|
||||
|
||||
```bash
|
||||
curl 'https://fuyao.aicubes.cn/api/a-share/prices/snapshot?thscodes=600519.SH' \
|
||||
-H "X-api-key: $HITHINK_FINANCE_API_KEY"
|
||||
```
|
||||
|
||||
## 5. 与小强现有架构的衔接建议
|
||||
|
||||
1. **回测数据**:用 marketdb 本地 DuckDB 替代/补充 Tushare 日线导入(现有 importer.py 继续保留,两者可交叉校验)
|
||||
2. **实时因子**:行情快照 API 做实时打分;涨跌停池/炸板池/连板天梯做情绪与事件因子
|
||||
3. **龙虎榜/异动/热榜**:事件驱动策略的增量数据源(Tushare 免费版这些覆盖弱)
|
||||
4. **财报/估值**:`financials` + `valuation` 命令批量拉全市场,喂多因子模型
|
||||
5. **全市场导出**:需要全量数据做研究时用 CLI Market Dumps,大结果落盘,避免上下文过载
|
||||
|
||||
## 6. 实测验证记录(2026-08-22)
|
||||
|
||||
```json
|
||||
{"thscode":"600519.SH","ticker":"600519","volume":3347231,"turnover":4278311000,
|
||||
"last_price":1272.83,"price_change":-18.67,"price_change_ratio_pct":-1.445606,
|
||||
"open_price":1291.5,"high_price":1291.5,"low_price":1272.01,"prev_price":1291.5}
|
||||
```
|
||||
|
||||
- ✅ auth login:`{"ok":true,"method":"api-key","configured":true}`
|
||||
- ✅ capabilities:symbol.search / market.snapshot / market.history / financials.* 等全量可用
|
||||
- ✅ 查询:贵州茅台 600519.SH 实时快照正常返回(周六休市=周五收盘数据)
|
||||
|
||||
## 7. 安全与合规提醒
|
||||
|
||||
- API Key 只存用户级凭据;不进代码、日志、公开配置、Git 仓库
|
||||
- Agent 交互时不得复述 Key;配置 Key 用 stdin 或环境变量
|
||||
- 数据仅用于研究/回测用途,遵守同花顺服务条款;不支持也不应伪造范围外数据(分钟K/tick/海外等)
|
||||
+245
-197
@@ -766,6 +766,179 @@ def import_daily_basic_by_date(
|
||||
logger.info(" 每日指标导入完成")
|
||||
|
||||
|
||||
# ============================================================
|
||||
# 4.1 导入资金流向 (moneyflow)
|
||||
# ============================================================
|
||||
|
||||
def import_moneyflow(
|
||||
ts_code: Optional[str] = None,
|
||||
trade_date: Optional[str] = None,
|
||||
start_date: Optional[str] = None,
|
||||
end_date: Optional[str] = None,
|
||||
conn=None,
|
||||
) -> int:
|
||||
"""
|
||||
导入个股资金流向 (moneyflow)
|
||||
Tushare: moneyflow
|
||||
|
||||
调用方式 (互斥,按优先级生效):
|
||||
1. 按单日全市场导入 (推荐,一次拉取全市场):
|
||||
import_moneyflow(trade_date="2026-08-14")
|
||||
对应示例: pro.moneyflow(trade_date='20260814')
|
||||
2. 按单只股票导入:
|
||||
import_moneyflow(ts_code="000001.SZ", start_date="2026-01-01", end_date="2026-08-14")
|
||||
对应示例: pro.moneyflow(ts_code='000001.SZ', start_date='20260101', end_date='20260814')
|
||||
3. 按交易日批量全市场导入请配合 import_moneyflow_by_date 使用
|
||||
|
||||
返回: 导入的记录数
|
||||
"""
|
||||
pro = get_ts_pro()
|
||||
own_conn = conn is None
|
||||
if own_conn:
|
||||
conn = get_pg_connection()
|
||||
|
||||
log_desc = ""
|
||||
try:
|
||||
# ---- 构建 moneyflow 请求参数 ----
|
||||
kwargs = {}
|
||||
if trade_date:
|
||||
# 按单日 (全市场)
|
||||
kwargs["trade_date"] = str(trade_date).replace("-", "")
|
||||
log_desc = f"交易日 {kwargs['trade_date']}"
|
||||
elif ts_code:
|
||||
# 按单只股票 (日期范围)
|
||||
if start_date is None:
|
||||
start_date = START_DATE
|
||||
if end_date is None:
|
||||
end_date = END_DATE
|
||||
kwargs["ts_code"] = ts_code
|
||||
kwargs["start_date"] = start_date.replace("-", "")
|
||||
kwargs["end_date"] = end_date.replace("-", "")
|
||||
log_desc = f"{ts_code} ({start_date} ~ {end_date})"
|
||||
else:
|
||||
logger.warning(
|
||||
" moneyflow: 请指定 trade_date (交易日, 如 '2026-08-14') 或 ts_code (股票代码)"
|
||||
)
|
||||
return 0
|
||||
|
||||
logger.info("=" * 60)
|
||||
logger.info(f"[4.1] 导入资金流向 (moneyflow): {log_desc}")
|
||||
|
||||
def fetch():
|
||||
return pro.moneyflow(**kwargs)
|
||||
|
||||
df = fetch_with_retry(fetch, max_retries=3)
|
||||
if df is None or df.empty:
|
||||
logger.warning(f" moneyflow ({log_desc}): 未获取到数据")
|
||||
return 0
|
||||
|
||||
df = normalize_columns(df)
|
||||
|
||||
# 转换日期列 (YYYYMMDD -> DATE)
|
||||
if "trade_date" in df.columns:
|
||||
df["trade_date"] = pd.to_datetime(df["trade_date"], format="%Y%m%d", errors="coerce")
|
||||
|
||||
# 数值列安全转换 (NaN -> None)
|
||||
numeric_cols = [
|
||||
"buy_sm_vol", "buy_sm_amount", "sell_sm_vol", "sell_sm_amount",
|
||||
"buy_md_vol", "buy_md_amount", "sell_md_vol", "sell_md_amount",
|
||||
"buy_lg_vol", "buy_lg_amount", "sell_lg_vol", "sell_lg_amount",
|
||||
"buy_elg_vol", "buy_elg_amount", "sell_elg_vol", "sell_elg_amount",
|
||||
"net_mf_vol", "net_mf_amount",
|
||||
]
|
||||
for col in numeric_cols:
|
||||
if col in df.columns:
|
||||
df[col] = pd.to_numeric(df[col], errors="coerce")
|
||||
|
||||
conflict_cols = ["ts_code", "trade_date"]
|
||||
return batch_insert("moneyflow", df, conn, conflict_cols)
|
||||
|
||||
finally:
|
||||
if own_conn:
|
||||
conn.close()
|
||||
|
||||
|
||||
def import_moneyflow_by_date(
|
||||
start_date: Optional[str] = None,
|
||||
end_date: Optional[str] = None,
|
||||
sleep_interval: float = 0.3,
|
||||
):
|
||||
"""
|
||||
按交易日批量导入资金流向 (全市场)
|
||||
Tushare moneyflow 接口可按交易日获取全市场数据,比较高效
|
||||
|
||||
用法:
|
||||
import_moneyflow_by_date(start_date="2010-01-01", end_date="2025-12-31")
|
||||
|
||||
返回: 失败的交易日列表
|
||||
"""
|
||||
if start_date is None:
|
||||
start_date = START_DATE
|
||||
if end_date is None:
|
||||
end_date = END_DATE
|
||||
|
||||
logger.info("=" * 60)
|
||||
logger.info(f"[4.1] 导入资金流向 (moneyflow): {start_date} ~ {end_date}")
|
||||
|
||||
# 获取交易日列表
|
||||
conn = get_pg_connection()
|
||||
try:
|
||||
cursor = conn.cursor()
|
||||
cursor.execute(
|
||||
"""
|
||||
SELECT DISTINCT cal_date FROM trade_cal
|
||||
WHERE is_open = 1
|
||||
AND cal_date >= %s AND cal_date <= %s
|
||||
ORDER BY cal_date
|
||||
""",
|
||||
(start_date, end_date),
|
||||
)
|
||||
trade_dates = [row[0].strftime("%Y%m%d") for row in cursor.fetchall()]
|
||||
cursor.close()
|
||||
finally:
|
||||
conn.close()
|
||||
|
||||
total = len(trade_dates)
|
||||
if total == 0:
|
||||
logger.warning(f" 日期范围 {start_date} ~ {end_date} 内无交易日")
|
||||
return []
|
||||
|
||||
logger.info(f" 共 {total} 个交易日")
|
||||
|
||||
conn = get_pg_connection()
|
||||
success_count = 0
|
||||
fail_list = []
|
||||
for i, td in enumerate(trade_dates, 1):
|
||||
try:
|
||||
pro = get_ts_pro()
|
||||
|
||||
def fetch_moneyflow():
|
||||
return pro.moneyflow(trade_date=td)
|
||||
|
||||
df = fetch_with_retry(fetch_moneyflow, max_retries=3)
|
||||
if df is not None and not df.empty:
|
||||
df = normalize_columns(df)
|
||||
if "trade_date" in df.columns:
|
||||
df["trade_date"] = pd.to_datetime(df["trade_date"], format="%Y%m%d", errors="coerce")
|
||||
conflict_cols = ["ts_code", "trade_date"]
|
||||
batch_insert("moneyflow", df, conn, conflict_cols)
|
||||
success_count += 1
|
||||
except Exception as e:
|
||||
logger.warning(f" [{td}] 导入失败: {e}")
|
||||
fail_list.append(td)
|
||||
conn.rollback()
|
||||
|
||||
if i % 20 == 0 or i == total:
|
||||
logger.info(f" 进度: {i}/{total} 成功={success_count} 失败={len(fail_list)}")
|
||||
time.sleep(sleep_interval)
|
||||
|
||||
conn.close()
|
||||
logger.info(f" 资金流向导入完成: 成功 {success_count}/{total}")
|
||||
if fail_list:
|
||||
logger.warning(f" 失败日期({len(fail_list)}): {fail_list[:20]}...")
|
||||
return fail_list
|
||||
|
||||
|
||||
# ============================================================
|
||||
# 5. 导入复权因子
|
||||
# ============================================================
|
||||
@@ -933,200 +1106,6 @@ def import_financial_statements(
|
||||
logger.info(" 财务数据导入完成")
|
||||
|
||||
|
||||
def import_cashflow(
|
||||
ts_code: Optional[str] = None,
|
||||
period: Optional[str] = None,
|
||||
start_date: Optional[str] = None,
|
||||
end_date: Optional[str] = None,
|
||||
conn=None,
|
||||
) -> int:
|
||||
"""
|
||||
导入现金流量表 (cashflow)
|
||||
Tushare: cashflow_vip (VIP 接口)
|
||||
|
||||
调用方式 (互斥,按优先级生效):
|
||||
1. 按报告期全市场导入 (推荐,调用次数最少,一次拉取全市场某报告期):
|
||||
import_cashflow(period="20181231")
|
||||
对应示例: df2 = pro.cashflow_vip(period='20181231', fields='')
|
||||
2. 按单只股票导入:
|
||||
import_cashflow(ts_code="000001.SZ", start_date="2010-01-01", end_date="2025-12-31")
|
||||
3. 股票批量导入请配合 import_cashflow_batch 使用
|
||||
|
||||
返回: 导入的记录数
|
||||
"""
|
||||
pro = get_ts_pro()
|
||||
own_conn = conn is None
|
||||
if own_conn:
|
||||
conn = get_pg_connection()
|
||||
|
||||
log_desc = ""
|
||||
try:
|
||||
# ---- 构建 cashflow_vip 请求参数 ----
|
||||
kwargs = {}
|
||||
if period:
|
||||
# 按报告期 (全市场) 导入,period 形如 20181231 或 2018-12-31
|
||||
kwargs["period"] = str(period).replace("-", "")
|
||||
log_desc = f"报告期 {kwargs['period']}"
|
||||
elif ts_code:
|
||||
# 按单只股票导入 (报告期范围)
|
||||
if start_date is None:
|
||||
start_date = START_DATE
|
||||
if end_date is None:
|
||||
end_date = END_DATE
|
||||
kwargs["ts_code"] = ts_code
|
||||
kwargs["start_date"] = start_date.replace("-", "")
|
||||
kwargs["end_date"] = end_date.replace("-", "")
|
||||
log_desc = f"{ts_code} ({start_date} ~ {end_date})"
|
||||
else:
|
||||
logger.warning(
|
||||
" cashflow: 请指定 period (报告期, 如 '20241231') 或 ts_code (股票代码)"
|
||||
)
|
||||
return 0
|
||||
|
||||
logger.info("=" * 60)
|
||||
logger.info(f"[6.1] 导入现金流量表 (cashflow_vip): {log_desc}")
|
||||
|
||||
def fetch():
|
||||
return pro.cashflow_vip(**kwargs)
|
||||
|
||||
df = fetch_with_retry(fetch, max_retries=3)
|
||||
if df is None or df.empty:
|
||||
logger.warning(f" cashflow ({log_desc}): 未获取到数据")
|
||||
return 0
|
||||
|
||||
df = normalize_columns(df)
|
||||
|
||||
# 转换日期列 (YYYYMMDD -> DATE)
|
||||
for col in ["ann_date", "f_ann_date", "end_date"]:
|
||||
if col in df.columns:
|
||||
df[col] = pd.to_datetime(df[col], format="%Y%m%d", errors="coerce")
|
||||
|
||||
conflict_cols = ["ts_code", "end_date", "report_type"]
|
||||
n = batch_insert("cashflow", df, conn, conflict_cols)
|
||||
return n
|
||||
|
||||
finally:
|
||||
if own_conn:
|
||||
conn.close()
|
||||
|
||||
|
||||
def import_cashflow_batch(
|
||||
stock_list: List[str],
|
||||
start_date: Optional[str] = None,
|
||||
end_date: Optional[str] = None,
|
||||
) -> int:
|
||||
"""
|
||||
按股票列表批量导入现金流量表 (cashflow)
|
||||
- 逐只股票调用 cashflow_vip (VIP 接口)
|
||||
- 适用于按股票维度补数据;全市场按报告期请用 import_cashflow(period=...)
|
||||
|
||||
返回: 累计导入的记录数
|
||||
"""
|
||||
if start_date is None:
|
||||
start_date = START_DATE
|
||||
if end_date is None:
|
||||
end_date = END_DATE
|
||||
|
||||
total = len(stock_list)
|
||||
logger.info("=" * 60)
|
||||
logger.info(
|
||||
f"[6.2] 批量导入现金流量表 (cashflow_vip): {start_date} ~ {end_date}, "
|
||||
f"共 {total} 只股票"
|
||||
)
|
||||
|
||||
conn = get_pg_connection()
|
||||
success = 0
|
||||
for i, ts_code in enumerate(stock_list, 1):
|
||||
try:
|
||||
n = import_cashflow(
|
||||
ts_code=ts_code,
|
||||
start_date=start_date,
|
||||
end_date=end_date,
|
||||
conn=conn,
|
||||
)
|
||||
success += n
|
||||
except Exception as e:
|
||||
logger.warning(f" [{ts_code}] 现金流量表导入失败: {e}")
|
||||
conn.rollback()
|
||||
|
||||
if i % 50 == 0 or i == total:
|
||||
logger.info(f" 进度: {i}/{total}, 累计导入 {success} 条")
|
||||
time.sleep(0.3)
|
||||
|
||||
conn.close()
|
||||
logger.info(f" 现金流量表批量导入完成, 共导入 {success} 条")
|
||||
return success
|
||||
|
||||
|
||||
def import_cashflow_initial(
|
||||
start_date: str = "2010-01-01",
|
||||
end_date: str = "2015-12-31",
|
||||
sleep_interval: float = 0.3,
|
||||
) -> int:
|
||||
"""
|
||||
首次批量导入现金流量表 (cashflow) - 按报告期全市场循环导入
|
||||
- 自动生成 start_date ~ end_date 范围内的所有季度报告期 (0331/0630/0930/1231)
|
||||
- 每个报告期调用一次 cashflow_vip(period=...) 一次性拉取全市场数据
|
||||
- 适用于首次初始化导入;后续增量/修补请用 import_cashflow(period=...) 或 import_cashflow_batch
|
||||
|
||||
示例:
|
||||
import_cashflow_initial(start_date="2010-01-01", end_date="2015-12-31")
|
||||
内部循环调用: pro.cashflow_vip(period='20100331'), pro.cashflow_vip(period='20100630'), ...,
|
||||
pro.cashflow_vip(period='20151231'),共 24 个报告期
|
||||
|
||||
返回: 累计导入的记录数
|
||||
"""
|
||||
# 解析年份范围 (支持 "2010-01-01" / "20100101" / "2010" 等格式)
|
||||
start_compact = str(start_date).replace("-", "")
|
||||
end_compact = str(end_date).replace("-", "")
|
||||
start_year = int(start_compact[:4])
|
||||
end_year = int(end_compact[:4])
|
||||
|
||||
# 自动生成所有季度报告期 (YYYYMMDD)
|
||||
periods = []
|
||||
for year in range(start_year, end_year + 1):
|
||||
for month_day in ["0331", "0630", "0930", "1231"]:
|
||||
period_str = f"{year}{month_day}"
|
||||
# 过滤掉首尾年份中超出日期范围的报告期
|
||||
if period_str < start_compact or period_str > end_compact:
|
||||
continue
|
||||
periods.append(period_str)
|
||||
|
||||
if not periods:
|
||||
logger.warning(f" 日期范围 {start_date} ~ {end_date} 内无报告期")
|
||||
return 0
|
||||
|
||||
total = len(periods)
|
||||
logger.info("=" * 60)
|
||||
logger.info(
|
||||
f"[6.3] 首次批量导入现金流量表 (cashflow_vip 按报告期): "
|
||||
f"{start_date} ~ {end_date}"
|
||||
)
|
||||
logger.info(f" 共 {total} 个报告期: {periods[0]} ~ {periods[-1]}")
|
||||
|
||||
conn = get_pg_connection()
|
||||
success = 0
|
||||
fail_list = []
|
||||
for i, period in enumerate(periods, 1):
|
||||
try:
|
||||
n = import_cashflow(period=period, conn=conn)
|
||||
success += n
|
||||
except Exception as e:
|
||||
logger.warning(f" [{period}] 现金流量表导入失败: {e}")
|
||||
fail_list.append(period)
|
||||
conn.rollback()
|
||||
|
||||
if i % 5 == 0 or i == total:
|
||||
logger.info(f" 进度: {i}/{total}, 累计导入 {success} 条")
|
||||
time.sleep(sleep_interval)
|
||||
|
||||
conn.close()
|
||||
logger.info(f" 首次现金流量表批量导入完成: 成功 {success} 条")
|
||||
if fail_list:
|
||||
logger.warning(f" 失败报告期({len(fail_list)}): {fail_list}")
|
||||
return success
|
||||
|
||||
|
||||
# ============================================================
|
||||
# 7. 导入指数日线行情
|
||||
# ============================================================
|
||||
@@ -1385,9 +1364,10 @@ def full_import(
|
||||
3. 交易日历
|
||||
4. 日线行情
|
||||
5. 每日指标(估值)
|
||||
6. 复权因子
|
||||
7. 财务数据 (可选)
|
||||
8. 指数日线行情
|
||||
6. 资金流向 (moneyflow)
|
||||
7. 复权因子
|
||||
8. 财务数据 (可选)
|
||||
9. 指数日线行情
|
||||
|
||||
参数:
|
||||
- start_date, end_date: 数据范围
|
||||
@@ -1430,6 +1410,9 @@ def full_import(
|
||||
# Step 4: 每日指标 (按日期导入)
|
||||
import_daily_basic_by_date(start_date, end_date)
|
||||
|
||||
# Step 4.1: 资金流向 (按交易日导入)
|
||||
import_moneyflow_by_date(start_date, end_date)
|
||||
|
||||
# Step 5: 复权因子
|
||||
import_adj_factor_batch(stock_codes, start_date, end_date)
|
||||
|
||||
@@ -1500,6 +1483,7 @@ def check_table_summary(conn=None):
|
||||
("trade_cal", (("trade_cal", "cal_date"),)),
|
||||
("daily", (("daily", "trade_date"),)),
|
||||
("daily_basic", (("daily_basic", "trade_date"),)),
|
||||
("moneyflow", (("moneyflow", "trade_date"),)),
|
||||
("adj_factor", (("adj_factor", "trade_date"),)),
|
||||
("income", (("income", "end_date"),)),
|
||||
("balancesheet", (("balancesheet", "end_date"),)),
|
||||
@@ -1582,6 +1566,70 @@ def get_missing_daily_dates(
|
||||
conn.close()
|
||||
|
||||
|
||||
def check_daily_coverage(
|
||||
start_date: Optional[str] = None,
|
||||
end_date: Optional[str] = None,
|
||||
conn=None,
|
||||
tolerance: float = 0.05,
|
||||
) -> List[Dict]:
|
||||
"""
|
||||
检查 daily 表每日覆盖度(以 daily_basic 为基准,单条 SQL 聚合)。
|
||||
|
||||
背景:resume_daily_by_date 只做「日期级」缺失检测(NOT EXISTS trade_date),
|
||||
若某交易日只有部分股票入库(如 2012-2013 沪市+创业板缺失),会误判为"已完整",
|
||||
导致覆盖度缺口永不补拉。本函数做「覆盖度级」校验:
|
||||
|
||||
daily 当日股票数 < daily_basic 当日股票数 × (1 - tolerance) → 判定覆盖不足
|
||||
|
||||
返回: [{"trade_date", "daily_cnt", "daily_basic_cnt", "coverage_pct"}, ...]
|
||||
"""
|
||||
if start_date is None:
|
||||
start_date = START_DATE
|
||||
if end_date is None:
|
||||
end_date = END_DATE
|
||||
|
||||
own_conn = conn is None
|
||||
if own_conn:
|
||||
conn = get_pg_connection()
|
||||
|
||||
try:
|
||||
cursor = conn.cursor()
|
||||
cursor.execute(
|
||||
"""
|
||||
SELECT db.trade_date,
|
||||
COALESCE(d.cnt, 0) AS daily_cnt,
|
||||
db.cnt AS db_cnt,
|
||||
ROUND(COALESCE(d.cnt, 0)::numeric / db.cnt * 100, 1) AS coverage_pct
|
||||
FROM (SELECT trade_date, COUNT(DISTINCT ts_code) AS cnt
|
||||
FROM daily_basic
|
||||
WHERE trade_date BETWEEN %s AND %s
|
||||
GROUP BY trade_date) db
|
||||
LEFT JOIN (SELECT trade_date, COUNT(DISTINCT ts_code) AS cnt
|
||||
FROM daily
|
||||
WHERE trade_date BETWEEN %s AND %s
|
||||
GROUP BY trade_date) d
|
||||
ON d.trade_date = db.trade_date
|
||||
WHERE COALESCE(d.cnt, 0) < db.cnt * (1 - %s)
|
||||
ORDER BY db.trade_date
|
||||
""",
|
||||
(start_date, end_date, start_date, end_date, tolerance),
|
||||
)
|
||||
rows = [
|
||||
{
|
||||
"trade_date": r[0].strftime("%Y-%m-%d") if hasattr(r[0], "strftime") else str(r[0])[:10],
|
||||
"daily_cnt": r[1],
|
||||
"daily_basic_cnt": r[2],
|
||||
"coverage_pct": float(r[3]),
|
||||
}
|
||||
for r in cursor.fetchall()
|
||||
]
|
||||
cursor.close()
|
||||
return rows
|
||||
finally:
|
||||
if own_conn:
|
||||
conn.close()
|
||||
|
||||
|
||||
def resume_daily_by_date(
|
||||
start_date: Optional[str] = None,
|
||||
end_date: Optional[str] = None,
|
||||
|
||||
@@ -18,7 +18,6 @@
|
||||
- 行情类表按 MAX(trade_date)+1 天 → 昨天 增量拉取
|
||||
- daily / daily_basic 走按交易日全市场模式(快);adj_factor 走按股票批量
|
||||
- 财务表按 MAX(end_date) 往前推 400 天 → 昨天(覆盖新公告的季度报告,UPSERT 幂等)
|
||||
- moneyflow 表暂未纳入(新版 importer 无对应导入函数,后续需要再补)
|
||||
"""
|
||||
import argparse
|
||||
import logging
|
||||
@@ -32,9 +31,11 @@ from importer import (
|
||||
import_stock_basic,
|
||||
import_daily_by_date,
|
||||
import_daily_basic_by_date,
|
||||
import_moneyflow_by_date,
|
||||
import_adj_factor_batch,
|
||||
import_index_daily,
|
||||
import_financial_statements,
|
||||
check_daily_coverage,
|
||||
)
|
||||
|
||||
logger = logging.getLogger("incremental")
|
||||
@@ -49,6 +50,7 @@ logger.setLevel(logging.INFO)
|
||||
TABLE_SPECS = {
|
||||
"daily": ("trade_date", ["ts_code", "trade_date"]),
|
||||
"daily_basic": ("trade_date", ["ts_code", "trade_date"]),
|
||||
"moneyflow": ("trade_date", ["ts_code", "trade_date"]),
|
||||
"adj_factor": ("trade_date", ["ts_code", "trade_date"]),
|
||||
"index_daily": ("trade_date", ["ts_code", "trade_date"]),
|
||||
"income": ("end_date", ["ts_code", "end_date", "report_type"]),
|
||||
@@ -105,6 +107,7 @@ def run_daily(end_date, dry_run, limit):
|
||||
for table, date_col in [
|
||||
("daily", "trade_date"),
|
||||
("daily_basic", "trade_date"),
|
||||
("moneyflow", "trade_date"),
|
||||
("adj_factor", "trade_date"),
|
||||
("index_daily", "trade_date"),
|
||||
]:
|
||||
@@ -129,6 +132,8 @@ def run_daily(end_date, dry_run, limit):
|
||||
import_daily_by_date(start, end_date)
|
||||
elif table == "daily_basic":
|
||||
import_daily_basic_by_date(start, end_date)
|
||||
elif table == "moneyflow":
|
||||
import_moneyflow_by_date(start, end_date)
|
||||
elif table == "adj_factor":
|
||||
import_adj_factor_batch(codes, start, end_date)
|
||||
elif table == "index_daily":
|
||||
@@ -136,6 +141,27 @@ def run_daily(end_date, dry_run, limit):
|
||||
except Exception as e:
|
||||
logger.error(f" {table} 增量导入失败: {e}")
|
||||
|
||||
# 5. 覆盖度校验(以 daily_basic 为基准,检查 daily 是否缺部分股票)
|
||||
# 防止"日期存在但覆盖不全"的缺口(如 2012-2013 沪市+创业板缺失)被增量逻辑跳过
|
||||
try:
|
||||
coverage_start = "2010-01-01" # 全历史检查(单条 SQL 聚合,开销小)
|
||||
logger.info(f"覆盖度校验: daily vs daily_basic ({coverage_start} ~ {end_date})")
|
||||
if dry_run:
|
||||
logger.info("[dry-run] 跳过覆盖度校验")
|
||||
else:
|
||||
partial = check_daily_coverage(coverage_start, end_date, conn=conn, tolerance=0.05)
|
||||
if partial:
|
||||
logger.warning(f"⚠ 发现 {len(partial)} 个交易日覆盖不足 (daily < daily_basic×95%):")
|
||||
for p in partial[:10]:
|
||||
logger.warning(f" {p['trade_date']}: daily={p['daily_cnt']} vs daily_basic={p['daily_basic_cnt']} ({p['coverage_pct']}%)")
|
||||
if len(partial) > 10:
|
||||
logger.warning(f" ... 其余 {len(partial)-10} 个交易日略")
|
||||
logger.warning(" 请运行 repair_daily_backfill.py 修复历史缺口,或检查近期导入是否被中断")
|
||||
else:
|
||||
logger.info(" ✓ 覆盖度正常,无缺失交易日")
|
||||
except Exception as e:
|
||||
logger.error(f" 覆盖度校验失败: {e}")
|
||||
|
||||
conn.close()
|
||||
logger.info("每日增量导入完成")
|
||||
|
||||
|
||||
@@ -0,0 +1,178 @@
|
||||
#!/usr/bin/env python3
|
||||
# -*- coding: utf-8 -*-
|
||||
"""
|
||||
修复 daily 表历史覆盖度缺口(2012-2013 沪市+创业板缺失)
|
||||
|
||||
背景:daily 表 2012-2013 年只有深市主板/中小板(2012 年 1383 只、2013 年 677 只),
|
||||
缺全部沪市 + 创业板。原因:当年全量导入沪市请求失败/中断,但深市成功,
|
||||
导致断点续传的"日期级"缺失检测(NOT EXISTS trade_date)认为该日已有数据,
|
||||
永不补拉。
|
||||
|
||||
修复策略(覆盖度级校验):
|
||||
1. 对指定日期范围,对比 daily 与 daily_basic 的当日去重股票数
|
||||
(daily_basic 同期数据完整,作为覆盖度基准)
|
||||
2. daily 当日股票数 < daily_basic 当日股票数 × (1 - tolerance) 的日期 → 判定为"部分缺失"
|
||||
3. 部分缺失的日期按全市场重新拉取(Tushare daily 按 trade_date 返回全市场)
|
||||
4. 用 INSERT ON CONFLICT DO NOTHING 幂等写入,可重复执行
|
||||
|
||||
用法:
|
||||
python3 repair_daily_backfill.py [--start 2012-01-01] [--end 2013-12-31] [--tolerance 0.05] [--dry-run]
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import argparse
|
||||
import logging
|
||||
from datetime import datetime
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
sys.path.insert(0, HERE)
|
||||
|
||||
# 显式加载 .env(config.py 会加载,但确保顺序正确)
|
||||
try:
|
||||
from dotenv import load_dotenv
|
||||
env_path = os.path.join(HERE, ".env")
|
||||
if os.path.exists(env_path):
|
||||
load_dotenv(env_path, override=False)
|
||||
except ImportError:
|
||||
pass
|
||||
|
||||
import importer
|
||||
from importer import (
|
||||
get_pg_connection, get_ts_pro, batch_insert,
|
||||
fetch_with_retry, _normalize_daily_df,
|
||||
)
|
||||
|
||||
logging.basicConfig(
|
||||
level=logging.INFO,
|
||||
format="%(asctime)s [%(levelname)s] %(message)s",
|
||||
handlers=[
|
||||
logging.FileHandler(os.path.join(HERE, "repair_daily_backfill.log"), encoding="utf-8"),
|
||||
logging.StreamHandler(),
|
||||
],
|
||||
)
|
||||
logger = logging.getLogger("repair_daily_backfill")
|
||||
|
||||
|
||||
def get_coverage_ratio(conn, trade_date: str) -> tuple:
|
||||
"""返回 (daily_stocks, daily_basic_stocks)。无参考数据时 daily_basic_stocks=None"""
|
||||
cur = conn.cursor()
|
||||
try:
|
||||
cur.execute("SELECT COUNT(DISTINCT ts_code) FROM daily WHERE trade_date=%s", (trade_date,))
|
||||
d_cnt = cur.fetchone()[0]
|
||||
cur.execute("SELECT COUNT(DISTINCT ts_code) FROM daily_basic WHERE trade_date=%s", (trade_date,))
|
||||
db_cnt = cur.fetchone()[0]
|
||||
return d_cnt, db_cnt
|
||||
finally:
|
||||
cur.close()
|
||||
|
||||
|
||||
def find_partial_dates(conn, start_date: str, end_date: str, tolerance: float) -> list:
|
||||
"""
|
||||
找出 daily 覆盖度不足的交易日(单条 SQL 聚合,避免逐日查询)。
|
||||
返回 [(trade_date, daily_cnt, daily_basic_cnt), ...]
|
||||
判据:daily_basic 有数据且 daily 股票数 < daily_basic × (1 - tolerance)
|
||||
"""
|
||||
cur = conn.cursor()
|
||||
try:
|
||||
cur.execute(
|
||||
"""
|
||||
SELECT db.trade_date,
|
||||
COALESCE(d.cnt, 0) AS daily_cnt,
|
||||
db.cnt AS db_cnt
|
||||
FROM (SELECT trade_date, COUNT(DISTINCT ts_code) AS cnt
|
||||
FROM daily_basic
|
||||
WHERE trade_date BETWEEN %s AND %s
|
||||
GROUP BY trade_date) db
|
||||
LEFT JOIN (SELECT trade_date, COUNT(DISTINCT ts_code) AS cnt
|
||||
FROM daily
|
||||
WHERE trade_date BETWEEN %s AND %s
|
||||
GROUP BY trade_date) d
|
||||
ON d.trade_date = db.trade_date
|
||||
WHERE COALESCE(d.cnt, 0) < db.cnt * (1 - %s)
|
||||
ORDER BY db.trade_date
|
||||
""",
|
||||
(start_date, end_date, start_date, end_date, tolerance),
|
||||
)
|
||||
partial = [(r[0].strftime("%Y-%m-%d"), r[1], r[2]) for r in cur.fetchall()]
|
||||
finally:
|
||||
cur.close()
|
||||
return partial
|
||||
|
||||
|
||||
def backfill_date(pro, conn, td_str: str) -> bool:
|
||||
"""拉取单个交易日全市场 daily 数据并写入。成功返回 True"""
|
||||
td_compact = td_str.replace("-", "")
|
||||
|
||||
def fetch():
|
||||
return pro.daily(trade_date=td_compact)
|
||||
|
||||
df = fetch_with_retry(fetch, max_retries=4, delay=3)
|
||||
if df is None or df.empty:
|
||||
logger.warning(f" [{td_str}] 返回空数据,跳过")
|
||||
return False
|
||||
|
||||
df = _normalize_daily_df(df)
|
||||
# 幂等写入:已存在的行跳过(ON CONFLICT DO NOTHING)
|
||||
batch_insert("daily", df, conn, ["ts_code", "trade_date"])
|
||||
return True
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="修复 daily 表历史覆盖度缺口")
|
||||
parser.add_argument("--start", default="2012-01-01", help="起始日期 YYYY-MM-DD")
|
||||
parser.add_argument("--end", default="2013-12-31", help="结束日期 YYYY-MM-DD")
|
||||
parser.add_argument("--tolerance", type=float, default=0.05,
|
||||
help="覆盖度容差,默认 0.05(daily 少于 daily_basic 的 95% 即判定缺失)")
|
||||
parser.add_argument("--dry-run", action="store_true", help="只扫描不导入")
|
||||
args = parser.parse_args()
|
||||
|
||||
conn = get_pg_connection()
|
||||
logger.info("=" * 60)
|
||||
logger.info(f"[覆盖度扫描] {args.start} ~ {args.end} (容差 {args.tolerance:.0%})")
|
||||
|
||||
partial = find_partial_dates(conn, args.start, args.end, args.tolerance)
|
||||
if not partial:
|
||||
logger.info(" 未发现覆盖度不足的交易日 ✅")
|
||||
conn.close()
|
||||
return
|
||||
|
||||
logger.info(f" 发现 {len(partial)} 个覆盖度不足的交易日:")
|
||||
# 按年份统计
|
||||
years = {}
|
||||
for td, d_cnt, db_cnt in partial:
|
||||
y = td[:4]
|
||||
years.setdefault(y, []).append((td, d_cnt, db_cnt))
|
||||
for y in sorted(years):
|
||||
lst = years[y]
|
||||
logger.info(f" {y} 年: {len(lst)} 个交易日 | 样例 {lst[0][0]}(daily={lst[0][1]}/db={lst[0][2]})")
|
||||
|
||||
if args.dry_run:
|
||||
logger.info("[dry-run] 不执行导入,以上为待补拉清单")
|
||||
conn.close()
|
||||
return
|
||||
|
||||
pro = get_ts_pro()
|
||||
success = 0
|
||||
fail_list = []
|
||||
for i, (td, d_cnt, db_cnt) in enumerate(partial, 1):
|
||||
ok = backfill_date(pro, conn, td)
|
||||
if ok:
|
||||
success += 1
|
||||
else:
|
||||
fail_list.append(td)
|
||||
if i % 20 == 0 or i == len(partial):
|
||||
logger.info(f" 进度: {i}/{len(partial)} 成功={success} 失败={len(fail_list)}")
|
||||
time.sleep(0.35) # 限流保护
|
||||
|
||||
conn.close()
|
||||
logger.info(f" 补拉完成: 成功 {success}/{len(partial)}")
|
||||
if fail_list:
|
||||
logger.warning(f" 失败日期({len(fail_list)}): {fail_list[:20]}...")
|
||||
# 写失败清单供重试
|
||||
with open(os.path.join(HERE, "repair_failed_dates.txt"), "w") as f:
|
||||
f.write("\n".join(fail_list))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -96,9 +96,6 @@
|
||||
" import_adj_factor,\n",
|
||||
" import_adj_factor_batch,\n",
|
||||
" import_financial_statements,\n",
|
||||
" import_cashflow,\n",
|
||||
" import_cashflow_batch,\n",
|
||||
" import_cashflow_initial,\n",
|
||||
" import_index_daily,\n",
|
||||
" get_all_stock_codes,\n",
|
||||
" get_stock_codes_from_db,\n",
|
||||
@@ -526,83 +523,13 @@
|
||||
")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### 7.1 单独导入现金流量表 (cashflow_vip VIP 接口)\n",
|
||||
"\n",
|
||||
"`import_cashflow` 使用 Tushare VIP 接口 `cashflow_vip`,支持两种方式:\n",
|
||||
"- 按报告期全市场导入: `import_cashflow(period=\"20181231\")`,对应示例 `df2 = pro.cashflow_vip(period='20181231', fields='')`\n",
|
||||
"- 按股票导入: `import_cashflow(ts_code=\"000001.SZ\", start_date=..., end_date=...)`,或批量 `import_cashflow_batch(stock_list, ...)`"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"#### 方式A2: 首次批量导入 (推荐初始化数据)\n",
|
||||
"\n",
|
||||
"`import_cashflow_initial` 自动生成日期范围内的所有季度报告期 (0331/0630/0930/1231),\n",
|
||||
"每个报告期调用一次 `cashflow_vip(period=...)` 一次拉取全市场数据。\n",
|
||||
"\n",
|
||||
"以下示例一次导入 **2010年1月1日 ~ 2015年12月31日** 共 24 个报告期的所有上市公司现金流量表数据。"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# 首次批量导入: 一次导入 2010-01-01 到 2015-12-31 的所有上市公司现金流量表\n",
|
||||
"# 自动循环 24 个季度报告期 (20100331 ~ 20151231),每个报告期拉取全市场数据\n",
|
||||
"import_cashflow_initial(\n",
|
||||
" start_date=\"2010-01-01\",\n",
|
||||
" end_date=\"2015-12-31\",\n",
|
||||
" sleep_interval=0.3,\n",
|
||||
")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# 方式A: 按报告期全市场导入 (推荐,一次拉取全市场某报告期数据,调用次数最少)\n",
|
||||
"# n = import_cashflow(period=\"20241231\") # 对应 pro.cashflow_vip(period='20241231', fields='')\n",
|
||||
"# print(f\"导入 20241231 报告期现金流量表: {n} 条\")\n",
|
||||
"\n",
|
||||
"# 方式B: 按股票列表批量导入 (适合补单只/部分股票数据)\n",
|
||||
"if 'stock_list' not in dir():\n",
|
||||
" stock_list = get_stock_codes_from_db()\n",
|
||||
"\n",
|
||||
"import_cashflow_batch(\n",
|
||||
" stock_list,\n",
|
||||
" start_date=\"2010-01-01\",\n",
|
||||
" end_date=\"2025-12-31\",\n",
|
||||
")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# 方式C: 导入单只股票的现金流量表\n",
|
||||
"# n = import_cashflow(ts_code=\"000001.SZ\", start_date=\"2010-01-01\", end_date=\"2025-12-31\")\n",
|
||||
"# print(f\"导入 000001.SZ 现金流量表: {n} 条\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# 验证财务数据\n",
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# 验证财务数据\n",
|
||||
"conn = get_pg_connection()\n",
|
||||
"cursor = conn.cursor()\n",
|
||||
"for table in [\"income\", \"balancesheet\", \"cashflow\", \"fina_indicator\"]:\n",
|
||||
@@ -793,6 +720,68 @@
|
||||
"display(df3[['ts_code', 'name', 'industry', 'roe', 'roa', 'eps', 'debt_to_assets']])\n",
|
||||
"conn.close()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"---\n",
|
||||
"## 数据完整性检验(覆盖度审计)\n",
|
||||
"\n",
|
||||
"> **背景**:2026-08-26 审计发现 daily 表 2012-2013 年存在覆盖度缺口(沪市+创业板整年缺失),\n",
|
||||
"> 而断点续传的\"日期级\"检测(`NOT EXISTS trade_date`)无法发现这种\"日期存在但覆盖不全\"的问题。\n",
|
||||
"> 本单元格做**覆盖度级**校验:以 `daily_basic` 当日股票数为基准,检查 `daily` 是否缺股。\n",
|
||||
">\n",
|
||||
"> 运行 `repair_daily_backfill.py` 可自动补拉覆盖不足的交易日。\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# ============================================================\n",
|
||||
"# 数据完整性检验:daily vs daily_basic 每日覆盖度审计\n",
|
||||
"# 判据:daily 当日股票数 < daily_basic 当日股票数 × 95% → 覆盖不足\n",
|
||||
"# ============================================================\n",
|
||||
"import importlib\n",
|
||||
"import importer\n",
|
||||
"importlib.reload(importer)\n",
|
||||
"from importer import get_pg_connection, check_daily_coverage, check_table_summary\n",
|
||||
"\n",
|
||||
"conn = get_pg_connection()\n",
|
||||
"\n",
|
||||
"# 1) 各表整体概览(行数 / 股票数 / 日期范围)\n",
|
||||
"print('=' * 60)\n",
|
||||
"print('各表整体概览:')\n",
|
||||
"print('=' * 60)\n",
|
||||
"check_table_summary(conn=conn)\n",
|
||||
"\n",
|
||||
"# 2) 全历史覆盖度审计(默认容差 5%)\n",
|
||||
"print('\\n' + '=' * 60)\n",
|
||||
"print('覆盖度审计: daily vs daily_basic (容差 5%)')\n",
|
||||
"print('=' * 60)\n",
|
||||
"from datetime import date\n",
|
||||
"end_date = date.today().strftime('%Y-%m-%d')\n",
|
||||
"partial = check_daily_coverage(start_date='2010-01-01', end_date=end_date, conn=conn, tolerance=0.05)\n",
|
||||
"\n",
|
||||
"if partial:\n",
|
||||
" print(f'⚠ 发现 {len(partial)} 个交易日覆盖不足:')\n",
|
||||
" # 按年份汇总\n",
|
||||
" from collections import Counter\n",
|
||||
" years = Counter(p['trade_date'][:4] for p in partial)\n",
|
||||
" for y in sorted(years):\n",
|
||||
" print(f' {y} 年: {years[y]} 个交易日覆盖不足')\n",
|
||||
" print('\\n样例(前10条):')\n",
|
||||
" for p in partial[:10]:\n",
|
||||
" print(f\" {p['trade_date']}: daily={p['daily_cnt']} vs daily_basic={p['daily_basic_cnt']} ({p['coverage_pct']}%)\")\n",
|
||||
" print('\\n→ 修复方法: 运行 python3 repair_daily_backfill.py 补拉')\n",
|
||||
"else:\n",
|
||||
" print('✅ 覆盖度正常,无缺失交易日')\n",
|
||||
"\n",
|
||||
"conn.close()\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
|
||||
Reference in New Issue
Block a user