Files
quant_engine/README.md
T
2026-10-04 11:42:23 +08:00

427 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# quant_engine
> 量化研究引擎 —— alpha 因子库 + 执行仿真 + 技术指标 + 数据适配 + 回测工具
**从 `research_results` 抽出的纯回测能力库**(v1.2.0 重构)。
## 角色
`quant_engine` 是 researchhub_workspace 的**引擎层**:
| 仓库 | 角色 |
|---|---|
| `quant_engine` | **纯研究核心**(alpha + execution + ledger + attribution + risk + metrics) |
| `research_results` | 业务集成(47 个 proj 调度 + 注册 + 平台对接) |
| `tushare2db_pro_aoge` | 数据层(行情 ELT) |
| `research_platform` | 展示层(FastAPI + Next.js) |
| `edb_data_core` | 数据层(经济数据) |
## 模块
- `alpha_factors` — 158 alpha 公式 + 24 基础算子(移植自 qlib alpha158)
- `factor_contracts` — `FactorDefinition` / `FactorSetRef` v1 纯计算合同、严格 PIT/availability 输入准入与显式 legacy 投影
- `execution` — A 股长仓执行仿真(成本/滑点/现金约束)+ 稀疏调仓/完整交易日 Ledger + 可投影成交与 NAV 审计;T+1、涨跌停、成交量与价差提供独立约束函数
- `indicators` — 50+ 技术指标(MACD / KDJ / 布林 / ATR / ADX / 等)
- `data_adapter` — 桥接 qtdb_pro 长表与新模块(rename / long-wide / 复权 / vwap 代理)
- `backtest` — weight-based 多日仿真(rebalance_table / compute_nav / compare_to_benchmark)
- `portfolio_construction` — 多期因子分数 → Top-K → 等权目标权重表
- `research_pipeline` — 因子日 → 下一真实交易日 → 显式执行价 → 日末估值 → 成本后绩效(防前视编排)
- `governed_pipeline` — 数据快照 → 因子版本 → 策略版本 → 回测运行 → 目标组合 → 风险决策 → Paper 订单意图;同时拥有输入/配置/重放血缘决定的 `BacktestRunRef`
- `artifact` — 版本化、确定性、存储中立的完整 research run 事实表,以及只映射现有表的 `BacktestEvidenceManifest`
- `portfolio_risk_contracts` — S3 证据闭合的 `PortfolioDecision` / `RiskAssessment` v1;独立复核 freshness、约束与 computation receipt,并复用既有标签安全风险分解
- `retrospective_*_contracts` — 未发布的显式 v2 回顾性合同:区分历史业务日期与实际可得/计算时间,保留 v1 和现有金融公式,不授予历史可得性、发布或执行权限;见 [v2 接口说明](docs/RETROSPECTIVE_COMPUTATION_V2.md)
- `attribution` — 基于实际成交后持仓的隔夜 / 日内 / 交易成本逐日收益归因与闭合审计
- `metrics` — 绝对绩效 + 严格日期对齐的 TE / IR / alpha / beta 基准相对绩效
- `strategy_research` / `strategy_optimizer` / `trade_pairing` / `strategy_artifact` — 七策略候选、下一日开盘执行、统一账本、FIFO双边成本、有界真实优化和完整策略报告工件;见 [策略合同](docs/strategy-research.md)
- `factor_diagnostics` — 候选0.1.0:完整键配对、逐日IC/RankIC、样本与未定义值、显式日历前瞻标签;见 [诊断合同](docs/factor-diagnostics.md)
- `factor_library` — 通用方法(turnover / winsorize / IC / OLS / jb_test)
- `portfolio_decomp` — 组合分解(risk_parity / mean_variance / 因子归因)
- `risk` — ndarray 低层风险公式 + 标签安全、可分组的 Euler 成分风险分解
- `perf_stats` — 详细绩效(与 metrics 并存)
- `logging` — 统一 logger(标准库 + 可选 loguru)
## 依赖
- 必需:numpy / pandas / scipy(标准量化栈)
- 可选:loguru(logback,标准库 logging 兜底)
**零重型依赖** —— 不引入 torch / lightgbm / hikyuu 等。
## 安装
```bash
cd quant_engine
pip install -e ".[dev]"
```
## 测试
```bash
pytest # 单元测试
pytest --cov=src # 覆盖率
mypy --strict src/ # 类型检查
ruff check src/ tests/ # lint
# 无网络、无数据库、无券商的架构烟测
uv run python -m quant_engine.governed_pipeline
```
## 使用
```python
from quant_engine.alpha_factors import alpha_001, alpha_005, ALPHA158_REGISTRY
from quant_engine.execution import (
ExecutionConfig, simulate_daily_ledger_with_audit,
simulate_multi_day_with_audit, simulate_with_daily_data,
)
from quant_engine.research_pipeline import (
run_factor_backtest_research, run_factor_execution_research,
)
from quant_engine.backtest import run_weight_backtest
from quant_engine.indicators import macd, bollinger, kdj
from quant_engine.data_adapter import (
long_to_wide, wide_to_long, rename_tushare_columns,
add_vwap_proxy, apply_adj_factor,
prepare_stock_series, prepare_execution_inputs,
load_qtdb_daily,
)
# 端到端:qtdb_pro 长表 → 适配 → alpha158 → execution
df = load_qtdb_daily(["000001.SZ"], "2024-01-01", with_adj=True)
close_prices, volumes = prepare_execution_inputs(df)
open_prices, _ = prepare_execution_inputs(df, price_col="open")
result = simulate_with_daily_data(close_prices, initial_cash=1_000_000.0)
# 已正确滞后的目标权重 → 现金约束执行 → 唯一来源的成交/拒绝/日末持仓/NAV
execution = simulate_multi_day_with_audit(
target_weights_history=[
("2024-01-02", {"000001.SZ": 1.0}),
("2024-01-03", {"000001.SZ": 1.0}),
],
price_history=[
("2024-01-02", {"000001.SZ": 10.0}),
("2024-01-03", {"000001.SZ": 10.5}),
],
initial_cash=1_000_000.0,
config=ExecutionConfig(),
)
print(execution.nav_series)
print(execution.daily_executions)
# 多期因子分数(必须是 point-in-time 数据)→ Top-K → 下一交易日 open 执行
factor_execution = run_factor_execution_research(
factor_scores,
top_k=20,
execution_prices=open_prices,
execution_price_field="open",
initial_cash=1_000_000.0,
)
# 推荐研究入口:同一交易日历上显式区分 open 成交和 close 估值。
# 因子日保持现金,下一交易日成交后的真实持仓才参与当日收盘收益。
factor_backtest = run_factor_backtest_research(
factor_scores,
top_k=20,
execution_prices=open_prices,
valuation_prices=close_prices,
execution_price_field="open",
valuation_price_field="close",
initial_cash=1_000_000.0,
config=ExecutionConfig(),
)
print(factor_backtest.nav)
print(factor_backtest.returns)
print(factor_backtest.stats())
print(factor_backtest.execution.ledger_frame)
print(factor_backtest.execution.trades_frame)
print(factor_backtest.position_weights) # 实际日末资产权重
print(factor_backtest.cash_weights)
# 所有分析都以实际成交后的 Ledger 为事实源,不直接使用目标权重伪造结果。
attribution = factor_backtest.return_attribution()
print(attribution.asset_contributions)
print(attribution.transaction_cost)
print(attribution.residual) # 应接近 0;否则说明贡献未闭合到账本收益
# benchmark_returns 必须与成本后 factor_backtest.returns 使用完全相同的日期索引。
print(factor_backtest.benchmark_stats(benchmark_returns))
# 下游稳定交付:显式提供代码版本、数据快照和时区,不在核心层写数据库。
from quant_engine.artifact import build_research_run_artifact
from quant_engine.data_adapter import prepare_asset_return_snapshot
from quant_engine.risk import estimate_covariance_snapshot
risk_date = factor_backtest.position_weights.index[-1].date()
market_snapshot = prepare_asset_return_snapshot(
qtdb_daily_long,
source="qtdb_pro.hq_daily",
source_snapshot_id="<upstream-ingestion-snapshot-id>",
adjustment="qfq",
)
risk_snapshot = estimate_covariance_snapshot(
market_snapshot.returns,
as_of_date=risk_date,
lookback_sessions=252,
min_observations=120,
data_snapshot_id=market_snapshot.data_snapshot_id,
)
artifact = build_research_run_artifact(
factor_backtest,
run_id="research-run-001",
strategy_id="alpha-top20",
strategy_name="Alpha Top 20",
strategy_version="1.0.0",
engine_version="1.2.0",
code_revision="<git-sha>",
data_snapshot_id=market_snapshot.data_snapshot_id,
calendar="CN-A",
timezone="Asia/Shanghai",
started_at="2026-08-21T10:00:00+08:00",
finished_at="2026-08-21T10:01:00+08:00",
parameters={"top_k": 20, "lag_sessions": 1},
benchmark_id="000300.SH",
benchmark_returns=benchmark_returns,
risk_snapshots={risk_date: risk_snapshot},
)
print(artifact.manifest())
# run_weight_backtest 是低层算子:只接受收益区间开始前已经生效的持仓权重。
# 不要把 signal-date 的 factor_scores/decision_weights 直接传给它。
backtest = run_weight_backtest(
weights=effective_holding_weights,
stock_returns=daily_returns,
initial_capital=1_000_000.0,
benchmark_nav=benchmark_nav,
)
print(factor_execution.schedule.signal_to_execution)
print(factor_execution.execution.daily_executions)
print(backtest.stats())
print(backtest.benchmark_report())
```
## 因子/特征合同 v1
`quant_engine.factor_contracts` 提供 `researchhub.factor-definition` 与
`researchhub.factor-set-ref` `1.0.0`。合同使用受限 canonical JSON:只接受 ASCII
lower-snake-case object key、UTF-8 string、bool/null 和 safe integer;小数参数必须用显式
canonical decimal string。定义、输入映射、上游证据、输出 schema/content 和 lineage 的任一
语义变化都会产生新 identity。
创建 `FactorSetRef` 必须提供完整且可重算 identity 的 `DatasetSnapshotEnvelope` 与
`DataFoundationEnvelope`,不能用 ID 字符串或布尔值代替资格证明。每个因子输入都要映射到一个
实际选中的 `StandardizedViewRef`,schema 必须同时匹配定义和 view;未消费、缺失、重复或跨
snapshot/Foundation/PIT 的 view 都会失败关闭。snapshot PIT 可以早于 Foundation/view PIT,
但始终满足 knowledge ≤ snapshot PIT ≤ Foundation/view/FactorSet PIT ≤ evaluation。
```python
from quant_engine.factor_contracts import (
DataFoundationEnvelope,
DatasetSnapshotEnvelope,
FactorDefinition,
FactorSetRef,
)
snapshot = DatasetSnapshotEnvelope.from_dict(dataset_snapshot_v1)
foundation = DataFoundationEnvelope.from_dict(data_foundation_v1)
# definition 必须是完整的 FactorDefinition;FactorSetRef.create 还要求显式 input bindings、
# view availability、output quality/coverage、canonical output bytes 和 immutable artifact ref。
factor_set = FactorSetRef.create(
definitions=(definition,),
dataset_snapshot=snapshot,
foundation=foundation,
**explicit_factor_set_evidence,
)
```
`availability_mode="as_available"` 声明 source/view 和计算产物在历史 evaluation 前实际可用;
`"retrospective_replay"` 保留历史 evaluation,但要求真实 publication/view creation、compute 和
artifact 时间位于之后,并固定 `historical_availability="not_established"`。两种模式都不会授予
decision、real-data、production、paper 或 live readiness。
旧 `governed_pipeline.FactorVersion` 的四字段构造器、`factor_id@version`、run/target/risk/order
identity 均保持不变。迁移只能通过 content-addressed `LegacyFactorBinding`,再显式调用
`bind_legacy_factor()` 或 `project_legacy_factor()`;后者是有损投影,不表示旧 digest 与新定义
digest 等价,也不会把旧 run 静默升级为新合同。
## 回测引用与证据合同 v1
`quant_engine.governed_pipeline.BacktestRunRef` 是合格回测运行身份的唯一权威。`run_id` 只由
已验收的 Dataset Snapshot / Data Foundation / `FactorSetRef` 身份、universe、日历与公司行动
祖先、策略、执行/成本模型、严格整数 seed、完整代码提交、环境锁、配置、时间和重放血缘决定;
它不包含任何输出摘要。重放必须绑定直接父运行、连续 attempt 和不变的
`replay_spec_digest`,输入漂移或血缘环会失败关闭。
`quant_engine.artifact.BacktestEvidenceManifest` 只摘要 `ResearchRunArtifact` 已有的九张事实表。
固定 `offline_research_v1` 映射为 `run`、`signal`、`fill`、`position_nav`、`performance`、
`attribution`、`risk_snapshot` 与 `replay`;每张表都保留列模式摘要、行数和内容摘要,空 risk
表也必须有稳定 schema。`manifest_id` 由完整 RunRef 与输出证据决定,因此结果变化不会反向改变
`run_id`。当前 artifact 不拥有订单或拒绝事实,所以此画像明确不声明 `order` / `rejection`。
旧 `BacktestRun` 只能通过 `build_legacy_backtest_evidence_manifest()` 显式映射为
`LEGACY_EXPLORATORY`;不能隐式提升为 `CONTRACT_QUALIFIED`。所有资格均只描述离线证据闭合,
不表示投资有效、组合获批、Paper、生产或实盘就绪。
## 绩效证据与方法论合同 v1
`quant_engine.artifact.PerformanceEvidenceV1` 在现有计算和事实表之外增加一层只读、内容寻址的
owner 证据。`build_performance_evidence()` 只接受同一运行的完整 `ResearchRunArtifact`、
`BacktestRunRef` 与 `CONTRACT_QUALIFIED BacktestEvidenceManifest`;它核对全部 artifact 表、
performance 表和唯一行摘要,并绑定 artifact、row 与严格对齐 benchmark series 的独立摘要。
生产 builder 不重算、填补、重命名或覆盖任何绩效值。
方法论固定为日简单收益、252 期年化、绝对指标年化无风险利率 `0.0`、benchmark 日无风险利率
`0.0`,以及 benchmark 存在时的 `exact_session_index`。相对指标使用封闭 availability:无基准为
`benchmark_absent`;active variance、benchmark variance 或 alpha 几何年化域不足时分别使用
对应 `not_estimable_*` 原因。benchmark 存在时 tracking error 始终必须是有限非负值;null 不会
被转成零。
```python
from quant_engine.artifact import build_performance_evidence
performance_evidence = build_performance_evidence(
artifact,
backtest_run_ref,
backtest_evidence_manifest,
)
canonical_bytes = performance_evidence.canonical_bytes()
```
该合同范围固定为 `offline_research_only`。它不授予排名、推荐、决策、发布、论文、Paper、生产、
实盘、交易或投资建议权限,也不包含原始参数、returns、NAV、benchmark series、表字节、存储
locator、URI 或凭证。
## 组合决策与风险评估合同 v1
`quant_engine.portfolio_risk_contracts` 是现有计算 owner 外围的薄合同层。创建
`PortfolioDecision` 必须同时提供完整 `BacktestRunRef`、嵌入同一 RunRef 的
`CONTRACT_QUALIFIED` 非 legacy `BacktestEvidenceManifest`、现有 `PortfolioTarget`、
`FreshnessPolicy`、`ConstraintSetV1` 与 `ComputationReceipt`。适配器会从权威输入独立重算
receipt 的 input/constraint/output digest、敞口、持仓数和 L1 turnover 残差;receipt 自报
成功、fallback 或放宽 tolerance 均不能替代复核。
```python
from quant_engine.portfolio_risk_contracts import (
ComputationReceipt,
ConstraintSetV1,
FreshnessPolicy,
assess_portfolio_risk,
build_portfolio_decision,
compute_portfolio_receipt_digests,
)
freshness = FreshnessPolicy(
max_manifest_age_seconds=3600,
max_covariance_age_days=5,
)
constraints = ConstraintSetV1(
gross_exposure_max=1.0,
single_asset_max=0.10,
position_count_max=20,
turnover_max=0.30,
)
# 生产者先形成公开 canonical digest;decision 构建时仍会独立重算。
expected = compute_portfolio_receipt_digests(
backtest_run_ref=run_ref,
manifest=evidence_manifest,
target=portfolio_target,
objective_name="long_only_allocation",
objective_version="1.0.0",
objective_digest=objective_digest,
model_name="factor_weighting",
model_version="1.0.0",
model_digest=model_digest,
expected_return_digest=expected_return_digest,
covariance_digest=covariance_digest,
scenario_digest=scenario_digest,
constraints=constraints,
freshness_policy=freshness,
prior_weights=prior_weights,
)
receipt = ComputationReceipt(
algorithm="factor_weighting",
algorithm_version="1.0.0",
implementation_digest=implementation_digest,
parameter_digest=parameter_digest,
input_digest=expected["input_digest"],
constraint_digest=expected["constraint_digest"],
output_digest=expected["output_digest"],
status="completed",
solver_required=False,
solver_name=None,
solver_version=None,
solver_config_digest=None,
iterations=None,
objective_value=None,
max_constraint_residual=expected["max_constraint_residual"],
tolerance=1e-12,
computed_at=computed_at,
)
decision = build_portfolio_decision(
backtest_run_ref=run_ref,
manifest=evidence_manifest,
target=portfolio_target,
objective_name="long_only_allocation",
objective_version="1.0.0",
objective_digest=objective_digest,
model_name="factor_weighting",
model_version="1.0.0",
model_digest=model_digest,
expected_return_digest=expected_return_digest,
covariance_digest=covariance_digest,
scenario_digest=scenario_digest,
constraints=constraints,
freshness_policy=freshness,
receipt=receipt,
computed_at=computed_at,
prior_weights=prior_weights,
)
assessment = assess_portfolio_risk(
portfolio_decision=decision,
backtest_run_ref=run_ref,
manifest=evidence_manifest,
covariance=covariance_snapshot,
risk_model_name="euler_volatility",
risk_model_version="1.0.0",
risk_model_digest=risk_model_digest,
)
```
`source_universe_digest` 保留 S3 研究 universe 身份,`portfolio_asset_set_digest` 只描述实际
目标资产标签;二者不会互相冒充成员证明。风险评估在任何数值计算前要求 covariance、target、
RunRef 的 dataset identity 三方一致,并且只调用一次现有 `labeled_component_risk()`。合同中的
`qualified` 仅表示 S4.1 计算证据闭合,不授予 maker-checker、发布、订单、Paper、生产或实盘权限。
## 治理垂直切片
`governed_pipeline` 不复制因子、回测、组合或执行算法,只编排现有能力并补充版本与风险契约。
调用方必须显式提供 `DatasetSnapshot`、`FactorVersion`、`StrategyVersion`、代码提交和
`RiskPolicy`。模块只会生成 `environment="paper"` 的订单意图,不连接数据库、数据供应商或
券商;风险决策为拒绝时,订单意图固定为空,直接调用创建函数也会失败关闭。
该切片对应 ResearchHub 架构的首个可执行验收链路:
```text
DatasetSnapshot → FactorVersion → StrategyVersion → BacktestRun
→ PortfolioTarget → RiskDecision → PaperOrderIntent
```
平台总架构、五仓职责和十二层能力映射仍以 `research_platform/docs/architecture/` 为权威;
本仓只拥有纯计算与离线模拟合同。
## 与 research_results 的关系
`research_results` 依赖 `quant_engine`(通过 re-export 保持向后兼容):
```python
# research_results/src/shared/alpha_factors.py 现在是:
from quant_engine.alpha_factors import * # re-export
```
47 个 proj 的 import 路径**暂时不变**(`from src.shared.alpha_factors import ...` 仍可用)——后续逐步迁移到 `from quant_engine.alpha_factors import ...`。