Compare commits

..
Author SHA1 Message Date
ao gong 4d84f6a26a Merge remote-tracking branch 'origin/main' into codex/research-artifact-contract-20260821 2026-08-25 22:39:54 +08:00
ao gong c12c9ba335 fix(artifact): enforce risk data lineage 2026-08-25 22:39:45 +08:00
ao gong 7c7f0c06a7 docs(handoff): record market lineage validation 2026-08-24 10:44:29 +08:00
ao gong 741d5f1ad3 feat(data): snapshot asset returns with stable lineage 2026-08-24 10:42:34 +08:00
ao gong cbb56d2f52 docs: record covariance estimator decision 2026-08-21 23:49:06 +08:00
ao gong 524fc73852 feat: estimate deterministic covariance snapshots 2026-08-21 23:47:29 +08:00
ao gong 8ef8e3208e test: define deterministic covariance estimator contract 2026-08-21 23:45:40 +08:00
ao gong dc67f0e982 docs: record reproducible risk artifact contract 2026-08-21 23:41:21 +08:00
ao gong 1a60fef6e1 feat: publish reproducible portfolio risk facts 2026-08-21 23:28:34 +08:00
ao gong a5dc04bf7e test: define reproducible risk snapshot contract 2026-08-21 23:24:41 +08:00
ao gong a3cefe364d feat: add deterministic trade and signal identity 2026-08-21 22:32:37 +08:00
ao gong d6c3614301 test: require deterministic trade identity 2026-08-21 22:32:14 +08:00
ao gong 8addce4a68 wip: hand off research artifact contract 2026-08-21 22:31:14 +08:00
ao gong 2071d508d0 test: cover sortino boundary contracts 2026-08-21 22:30:28 +08:00
ao gong 0568cabda1 feat: include signal facts in research artifact 2026-08-21 22:30:01 +08:00
ao gong b2f3da9d66 test: require signal facts in research artifact 2026-08-21 22:29:32 +08:00
ao gong cfa5bed188 feat: add deterministic research run artifact 2026-08-21 22:29:10 +08:00
ao gong a9465e6479 test: define versioned research artifact contract 2026-08-21 22:26:51 +08:00
ao gong 55eeff3951 docs: add realized ledger weights to handoff 2026-08-21 22:22:00 +08:00
ao gong 3b1ad07c69 feat: expose realized ledger position weights 2026-08-21 22:21:32 +08:00
ao gong 1b5b353098 test: define realized ledger weight projection contract 2026-08-21 22:21:03 +08:00
ao gong 2bb9f52080 wip: hand off ledger-backed attribution 2026-08-21 22:19:20 +08:00
ao gong 76bb5494a2 docs: record lightweight attribution design references 2026-08-21 22:18:24 +08:00
ao gong f7ad82534a feat: add label-safe component risk decomposition 2026-08-21 22:17:05 +08:00
ao gong 35a52d781e test: define labeled component risk contract 2026-08-21 22:16:24 +08:00
ao gong f14ab464f7 feat: add strict benchmark-relative performance metrics 2026-08-21 22:15:47 +08:00
ao gong de2f9494fc test: define benchmark-relative performance contract 2026-08-21 22:15:05 +08:00
ao gong 19fe22b01a feat: add ledger-backed daily return attribution 2026-08-21 22:14:18 +08:00
ao gong 212351e984 test: define post-execution return attribution contract 2026-08-21 22:13:04 +08:00
ao gong 5fb4b85cc3 wip: hand off stacked daily ledger 2026-08-21 22:06:51 +08:00
ao gong a421278527 fix: align ledger performance with signal window 2026-08-21 22:05:28 +08:00
ao gong c774a4546a test: exclude factor warmup from performance window 2026-08-21 22:04:52 +08:00
ao gong 022b87fdac feat: expose platform-neutral ledger projection 2026-08-21 22:03:21 +08:00
ao gong 1b39c53f18 test: define daily ledger projection contract 2026-08-21 22:02:58 +08:00
ao gong b794ab2e8f feat: add post-execution daily ledger 2026-08-21 22:01:35 +08:00
ao gong a44e2d306a test: define post-execution daily ledger contract 2026-08-21 21:58:01 +08:00
ao gong 15b283bdf9 docs: distinguish signal execution and holding times
CI / lite (pull_request) Successful in 4s
2026-08-21 21:49:07 +08:00
ao gong f9b7f2ab1a feat: schedule factor weights for next-session execution 2026-08-21 21:47:53 +08:00
ao gong 213aa88deb test: require execution-price input snapshot 2026-08-21 21:46:47 +08:00
ao gong b7f77e2a6c test: define lagged factor-to-execution contract 2026-08-21 21:46:00 +08:00
ao gong c9fb5978f0 docs: clarify execution audit result contract
CI / lite (pull_request) Successful in 3s
2026-08-21 21:39:12 +08:00
ao gong b2af2a10a6 docs: document auditable execution workflow 2026-08-21 21:38:08 +08:00
ao gong da2ca51ff7 fix: enforce cash-backed long-only rebalancing 2026-08-21 21:37:02 +08:00
ao gong ee22d1eb9d test: reproduce execution cash and long-only violations 2026-08-21 21:35:47 +08:00
ao gong 17b680604d fix: derive multi-day trades from target-weight deltas 2026-08-21 21:35:15 +08:00
ao gong c51d803daf test: reproduce multi-day execution audit gaps 2026-08-21 21:31:52 +08:00
ao gong 5578851d85 fix: enforce chronological rebalance scores
CI / lite (pull_request) Canceled after 0s
2026-08-21 21:25:27 +08:00
ao gong b4f7b74c04 test: reject unsorted rebalance dates 2026-08-21 21:25:11 +08:00
ao gong 0a236622f3 docs: add factor-to-backtest portfolio workflow 2026-08-21 21:24:41 +08:00
ao gong 88af157ee7 fix: validate unique rebalance dates 2026-08-21 21:24:17 +08:00
ao gong a1130ea43b test: reject ambiguous duplicate rebalance dates 2026-08-21 21:24:03 +08:00
ao gong 96e8f1ad62 feat: build target weights from factor scores 2026-08-21 21:23:25 +08:00
ao gong ecf6b4e4bf test: add red factor-to-weights portfolio contract 2026-08-21 21:22:36 +08:00
ao gong 979166ad16 docs: document unified backtest result workflow
CI / lite (pull_request) Successful in 3s
2026-08-21 21:11:52 +08:00
ao gong cdf41edf1f feat: add unified weight backtest result facade 2026-08-21 21:11:06 +08:00
ao gong 585a635797 test: add red contract for unified backtest result 2026-08-21 21:10:33 +08:00
ao gong 2ff0630973 refactor: clarify quant core validation internals
CI / lite (pull_request) Successful in 3s
2026-08-21 21:00:19 +08:00
ao gong bae4dedf70 test: satisfy metric contract lint 2026-08-21 20:59:27 +08:00
ao gong e359792ef5 fix: harden quant core calculation boundaries 2026-08-21 20:58:30 +08:00
ao gong 0c3a375b1e test: add red contracts for quant core boundaries 2026-08-21 20:57:49 +08:00
5 changed files with 1 additions and 2733 deletions
-21
View File
@@ -25,7 +25,6 @@
- `backtest` — weight-based 多日仿真(rebalance_table / compute_nav / compare_to_benchmark) - `backtest` — weight-based 多日仿真(rebalance_table / compute_nav / compare_to_benchmark)
- `portfolio_construction` — 多期因子分数 → Top-K → 等权目标权重表 - `portfolio_construction` — 多期因子分数 → Top-K → 等权目标权重表
- `research_pipeline` — 因子日 → 下一真实交易日 → 显式执行价 → 日末估值 → 成本后绩效(防前视编排) - `research_pipeline` — 因子日 → 下一真实交易日 → 显式执行价 → 日末估值 → 成本后绩效(防前视编排)
- `governed_pipeline` — 数据快照 → 因子版本 → 策略版本 → 回测运行 → 目标组合 → 风险决策 → Paper 订单意图;全链路带确定性 ID,风险拒绝时禁止生成订单意图
- `artifact` — 版本化、确定性、存储中立的完整 research run 事实表与 manifest - `artifact` — 版本化、确定性、存储中立的完整 research run 事实表与 manifest
- `attribution` — 基于实际成交后持仓的隔夜 / 日内 / 交易成本逐日收益归因与闭合审计 - `attribution` — 基于实际成交后持仓的隔夜 / 日内 / 交易成本逐日收益归因与闭合审计
- `metrics` — 绝对绩效 + 严格日期对齐的 TE / IR / alpha / beta 基准相对绩效 - `metrics` — 绝对绩效 + 严格日期对齐的 TE / IR / alpha / beta 基准相对绩效
@@ -56,9 +55,6 @@ pytest # 单元测试
pytest --cov=src # 覆盖率 pytest --cov=src # 覆盖率
mypy --strict src/ # 类型检查 mypy --strict src/ # 类型检查
ruff check src/ tests/ # lint ruff check src/ tests/ # lint
# 无网络、无数据库、无券商的架构烟测
uv run python -m quant_engine.governed_pipeline
``` ```
## 使用 ## 使用
@@ -195,23 +191,6 @@ print(backtest.stats())
print(backtest.benchmark_report()) print(backtest.benchmark_report())
``` ```
## 治理垂直切片
`governed_pipeline` 不复制因子、回测、组合或执行算法,只编排现有能力并补充版本与风险契约。
调用方必须显式提供 `DatasetSnapshot`、`FactorVersion`、`StrategyVersion`、代码提交和
`RiskPolicy`。模块只会生成 `environment="paper"` 的订单意图,不连接数据库、数据供应商或
券商;风险决策为拒绝时,订单意图固定为空,直接调用创建函数也会失败关闭。
该切片对应 ResearchHub 架构的首个可执行验收链路:
```text
DatasetSnapshot → FactorVersion → StrategyVersion → BacktestRun
→ PortfolioTarget → RiskDecision → PaperOrderIntent
```
平台总架构、五仓职责和十二层能力映射仍以 `research_platform/docs/architecture/` 为权威;
本仓只拥有纯计算与离线模拟合同。
## 与 research_results 的关系 ## 与 research_results 的关系
`research_results` 依赖 `quant_engine`(通过 re-export 保持向后兼容): `research_results` 依赖 `quant_engine`(通过 re-export 保持向后兼容):
+1 -927
View File
@@ -15,9 +15,7 @@ v1.2.0 Phase 0:5 个基础算子 + 5 个 alpha 公式(alpha001–alpha005)
from __future__ import annotations from __future__ import annotations
from collections.abc import Callable, Mapping from typing import Any
from types import MappingProxyType
from typing import Any, cast
import numpy as np import numpy as np
import pandas as pd import pandas as pd
@@ -211,367 +209,6 @@ def indneutralize(series: pd.Series, groups: pd.Series) -> pd.Series:
return series - series.groupby(groups).transform("mean") return series - series.groupby(groups).transform("mean")
# ── Phase 1 operator contract ──────────────────────────
# This is deliberately a small, stable surface for downstream research
# orchestration. The full alpha158 formula catalogue can continue to grow,
# while callers use one validated dispatch entry point for the first ten
# deterministic building blocks.
ALPHA158_PHASE1_MAX_WINDOW = 252
ALPHA158_PHASE1_OPERATOR_SPECS: dict[str, dict[str, Any]] = {
"rank": {
"name": "rank",
"formula": "rank(series)",
"inputs": ["series"],
"windowed": False,
},
"delta": {
"name": "delta",
"formula": "delta(series, window)",
"inputs": ["series"],
"windowed": True,
},
"ts_mean": {
"name": "ts_mean",
"formula": "ts_mean(series, window)",
"inputs": ["series"],
"windowed": True,
},
"ts_std": {
"name": "ts_std",
"formula": "ts_std(series, window)",
"inputs": ["series"],
"windowed": True,
},
"ts_rank": {
"name": "ts_rank",
"formula": "ts_rank(series, window)",
"inputs": ["series"],
"windowed": True,
},
"correlation": {
"name": "correlation",
"formula": "correlation(series, secondary, window)",
"inputs": ["series", "secondary"],
"windowed": True,
},
"ts_min": {
"name": "ts_min",
"formula": "ts_min(series, window)",
"inputs": ["series"],
"windowed": True,
},
"ts_max": {
"name": "ts_max",
"formula": "ts_max(series, window)",
"inputs": ["series"],
"windowed": True,
},
"ts_sum": {
"name": "ts_sum",
"formula": "ts_sum(series, window)",
"inputs": ["series"],
"windowed": True,
},
"decay_linear": {
"name": "decay_linear",
"formula": "decay_linear(series, window)",
"inputs": ["series"],
"windowed": True,
},
}
_PHASE1_OPERATOR_FUNCTIONS: dict[str, Callable[..., pd.Series]] = {
"rank": rank,
"delta": delta,
"ts_mean": ts_mean,
"ts_std": ts_std,
"ts_rank": ts_rank,
"correlation": correlation,
"ts_min": ts_min,
"ts_max": ts_max,
"ts_sum": ts_sum,
"decay_linear": decay_linear,
}
def list_phase1_operators() -> tuple[str, ...]:
"""Return the deterministic Phase 1 operator names in stable order."""
return tuple(ALPHA158_PHASE1_OPERATOR_SPECS)
def evaluate_phase1_operator(
name: str,
series: pd.Series,
secondary: pd.Series | None = None,
*,
window: int | None = None,
) -> pd.Series:
"""Evaluate one of the ten Phase 1 operators with a validated contract.
``window`` is required for time-series operators and forbidden for the
cross-sectional ``rank`` operator. Binary ``correlation`` also requires
a same-index secondary series so that callers cannot silently introduce
alignment-dependent results.
"""
if name not in ALPHA158_PHASE1_OPERATOR_SPECS:
raise KeyError(f"operator {name!r} not registered")
is_windowed = bool(ALPHA158_PHASE1_OPERATOR_SPECS[name]["windowed"])
if is_windowed:
if isinstance(window, bool) or not isinstance(window, int) or window <= 0:
raise ValueError(f"window must be a positive integer for {name}")
if window > ALPHA158_PHASE1_MAX_WINDOW:
raise ValueError(
f"window exceeds maximum supported value {ALPHA158_PHASE1_MAX_WINDOW} for {name}"
)
if not is_windowed and window is not None:
raise ValueError(f"window is not supported for {name}")
if name == "correlation":
if secondary is None:
raise ValueError("secondary is required for correlation")
if not series.index.equals(secondary.index):
raise ValueError("secondary index must align with series")
return correlation(series, secondary, window) # type: ignore[arg-type]
if secondary is not None:
raise ValueError(f"secondary is not supported for {name}")
operator = _PHASE1_OPERATOR_FUNCTIONS[name]
if name == "rank":
return operator(series)
return operator(series, window)
# ── Phase 2 cumulative operator contract ──────────────────────────────
# Phase 2 is cumulative: downstream callers can upgrade to one dispatch
# surface covering every existing alpha158 building block, while Phase 1
# names, metadata, ordering, and evaluation remain unchanged.
ALPHA158_PHASE2_MAX_WINDOW = ALPHA158_PHASE1_MAX_WINDOW
ALPHA158_PHASE2_OPERATOR_SPECS: dict[str, dict[str, Any]] = {
name: {
**spec,
"parameters": ["window"] if bool(spec["windowed"]) else [],
}
for name, spec in ALPHA158_PHASE1_OPERATOR_SPECS.items()
}
ALPHA158_PHASE2_OPERATOR_SPECS.update(
{
"ts_argmin": {
"name": "ts_argmin",
"formula": "ts_argmin(series, window)",
"inputs": ["series"],
"parameters": ["window"],
"windowed": True,
},
"ts_argmax": {
"name": "ts_argmax",
"formula": "ts_argmax(series, window)",
"inputs": ["series"],
"parameters": ["window"],
"windowed": True,
},
"product": {
"name": "product",
"formula": "product(series, window)",
"inputs": ["series"],
"parameters": ["window"],
"windowed": True,
},
"returns": {
"name": "returns",
"formula": "returns(series)",
"inputs": ["series"],
"parameters": [],
"windowed": False,
},
"scale": {
"name": "scale",
"formula": "scale(series)",
"inputs": ["series"],
"parameters": [],
"windowed": False,
},
"signed_power": {
"name": "signed_power",
"formula": "signed_power(series, exponent)",
"inputs": ["series"],
"parameters": ["exponent"],
"windowed": False,
},
"stddev": {
"name": "stddev",
"formula": "stddev(series, window)",
"inputs": ["series"],
"parameters": ["window"],
"windowed": True,
},
"covariance": {
"name": "covariance",
"formula": "covariance(series, secondary, window)",
"inputs": ["series", "secondary"],
"parameters": ["window"],
"windowed": True,
},
"log": {
"name": "log",
"formula": "log(series)",
"inputs": ["series"],
"parameters": [],
"windowed": False,
},
"abs_series": {
"name": "abs_series",
"formula": "abs_series(series)",
"inputs": ["series"],
"parameters": [],
"windowed": False,
},
"sign": {
"name": "sign",
"formula": "sign(series)",
"inputs": ["series"],
"parameters": [],
"windowed": False,
},
"max_pair": {
"name": "max_pair",
"formula": "max_pair(series, secondary)",
"inputs": ["series", "secondary"],
"parameters": [],
"windowed": False,
},
"min_pair": {
"name": "min_pair",
"formula": "min_pair(series, secondary)",
"inputs": ["series", "secondary"],
"parameters": [],
"windowed": False,
},
"indneutralize": {
"name": "indneutralize",
"formula": "indneutralize(series, groups)",
"inputs": ["series", "groups"],
"parameters": [],
"windowed": False,
},
}
)
_PHASE2_OPERATOR_FUNCTIONS: dict[str, Callable[..., pd.Series]] = {
**_PHASE1_OPERATOR_FUNCTIONS,
"ts_argmin": ts_argmin,
"ts_argmax": ts_argmax,
"product": product,
"returns": returns,
"scale": scale,
"signed_power": signed_power,
"stddev": stddev,
"covariance": covariance,
"log": log,
"abs_series": abs_series,
"sign": sign,
"max_pair": max_pair,
"min_pair": min_pair,
"indneutralize": indneutralize,
}
_PHASE2_WINDOWED_OPERATORS = frozenset(
name for name, spec in ALPHA158_PHASE2_OPERATOR_SPECS.items() if bool(spec["windowed"])
)
_PHASE2_BINARY_OPERATORS = frozenset({"correlation", "covariance", "max_pair", "min_pair"})
def list_phase2_operators() -> tuple[str, ...]:
"""Return all Phase 2 operator names in stable cumulative order."""
return tuple(ALPHA158_PHASE2_OPERATOR_SPECS)
def _validate_phase2_window(name: str, window: int | None) -> int:
if isinstance(window, bool) or not isinstance(window, int) or window <= 0:
raise ValueError(f"window must be a positive integer for {name}")
if window > ALPHA158_PHASE2_MAX_WINDOW:
raise ValueError(
f"window exceeds maximum supported value {ALPHA158_PHASE2_MAX_WINDOW} for {name}"
)
return window
def evaluate_phase2_operator(
name: str,
series: pd.Series,
secondary: pd.Series | None = None,
*,
window: int | None = None,
exponent: float | None = None,
groups: pd.Series | None = None,
) -> pd.Series:
"""Evaluate any existing alpha158 building block through a strict contract.
Phase 2 rejects implicit alignment, missing required arguments, unused
arguments, unbounded windows, and non-finite exponents before dispatch.
"""
if name not in ALPHA158_PHASE2_OPERATOR_SPECS:
raise KeyError(f"operator {name!r} not registered")
if not isinstance(series, pd.Series):
raise TypeError("series must be a pandas Series")
validated_window: int | None = None
if name in _PHASE2_WINDOWED_OPERATORS:
validated_window = _validate_phase2_window(name, window)
elif window is not None:
raise ValueError(f"window is not supported for {name}")
if name in _PHASE2_BINARY_OPERATORS:
if secondary is None:
raise ValueError(f"secondary is required for {name}")
if not isinstance(secondary, pd.Series):
raise TypeError("secondary must be a pandas Series")
if not series.index.equals(secondary.index):
raise ValueError("secondary index must align with series")
elif secondary is not None:
raise ValueError(f"secondary is not supported for {name}")
validated_exponent: float | None = None
if name == "signed_power":
if (
isinstance(exponent, bool)
or not isinstance(exponent, (int, float))
or not np.isfinite(exponent)
):
raise ValueError("exponent must be a finite number for signed_power")
validated_exponent = float(exponent)
elif exponent is not None:
raise ValueError(f"exponent is not supported for {name}")
if name == "indneutralize":
if groups is None:
raise ValueError("groups is required for indneutralize")
if not isinstance(groups, pd.Series):
raise TypeError("groups must be a pandas Series")
if not series.index.equals(groups.index):
raise ValueError("groups index must align with series")
elif groups is not None:
raise ValueError(f"groups is not supported for {name}")
operator = _PHASE2_OPERATOR_FUNCTIONS[name]
if name == "signed_power":
return operator(series, validated_exponent)
if name == "indneutralize":
return operator(series, groups)
if name in {"correlation", "covariance"}:
return operator(series, secondary, validated_window)
if name in {"max_pair", "min_pair"}:
return operator(series, secondary)
if validated_window is not None:
return operator(series, validated_window)
return operator(series)
# ── 组合算子(alpha158 公式样本) ───────────────────────── # ── 组合算子(alpha158 公式样本) ─────────────────────────
@@ -3093,541 +2730,6 @@ def parse_alpha_formula(formula_str: str) -> dict[str, Any]:
return parsed return parsed
# ── Phase 3 formula contract: frozen alpha001-alpha050 surface ──────────────
# Formula functions remain the implementation source of truth. This contract
# freezes their callable surface separately from formula dependencies so that
# historical compatibility-only arguments remain explicit without rewriting
# formulas or changing direct-call APIs.
ALPHA158_PHASE3_FORMULA_CONTRACT_VERSION = "1.0.0"
ALPHA158_PHASE3_FORMULA_CATALOG_SHA256 = (
"9a3360e5ee77a85a35d3c2fdab1efaa531fa0c263a2cb2bb5b965a1fb96fe1bd"
)
_PHASE3_FORMULA_FUNCTIONS: dict[str, Callable[..., pd.Series]] = {
"alpha_001": alpha_001,
"alpha_002": alpha_002,
"alpha_003": alpha_003,
"alpha_004": alpha_004,
"alpha_005": alpha_005,
"alpha_006": alpha_006,
"alpha_007": alpha_007,
"alpha_008": alpha_008,
"alpha_009": alpha_009,
"alpha_010": alpha_010,
"alpha_011": alpha_011,
"alpha_012": alpha_012,
"alpha_013": alpha_013,
"alpha_014": alpha_014,
"alpha_015": alpha_015,
"alpha_016": alpha_016,
"alpha_017": alpha_017,
"alpha_018": alpha_018,
"alpha_019": alpha_019,
"alpha_020": alpha_020,
"alpha_021": alpha_021,
"alpha_022": alpha_022,
"alpha_023": alpha_023,
"alpha_024": alpha_024,
"alpha_025": alpha_025,
"alpha_026": alpha_026,
"alpha_027": alpha_027,
"alpha_028": alpha_028,
"alpha_029": alpha_029,
"alpha_030": alpha_030,
"alpha_031": alpha_031,
"alpha_032": alpha_032,
"alpha_033": alpha_033,
"alpha_034": alpha_034,
"alpha_035": alpha_035,
"alpha_036": alpha_036,
"alpha_037": alpha_037,
"alpha_038": alpha_038,
"alpha_039": alpha_039,
"alpha_040": alpha_040,
"alpha_041": alpha_041,
"alpha_042": alpha_042,
"alpha_043": alpha_043,
"alpha_044": alpha_044,
"alpha_045": alpha_045,
"alpha_046": alpha_046,
"alpha_047": alpha_047,
"alpha_048": alpha_048,
"alpha_049": alpha_049,
"alpha_050": alpha_050,
}
_PHASE3_FORMULA_INPUT_OVERRIDES: dict[str, list[str]] = {
"alpha_011": ["close", "high", "low"],
"alpha_035": ["volume"],
"alpha_036": ["close"],
"alpha_040": ["high", "low"],
"alpha_042": ["close"],
"alpha_043": ["volume"],
}
_PHASE3_INPUT_CATEGORIES = {
1: "single",
2: "pair",
3: "triple",
4: "quadruple",
}
def _phase3_call_inputs(function: Callable[..., pd.Series]) -> list[str]:
import inspect
parameters = list(inspect.signature(function).parameters.values())
if any(
parameter.kind is not inspect.Parameter.POSITIONAL_OR_KEYWORD
or parameter.default is not inspect.Parameter.empty
for parameter in parameters
):
raise RuntimeError(f"unsupported formula signature for {function.__name__}")
return ["open" if parameter.name == "open_" else parameter.name for parameter in parameters]
def _phase3_string_list(meta: dict[str, Any], field: str, alpha_id: str) -> list[str]:
value = meta[field]
if not isinstance(value, list) or not all(isinstance(item, str) for item in value):
raise RuntimeError(f"{field} must be a list of strings for {alpha_id}")
return list(value)
def _build_phase3_formula_specs() -> dict[str, dict[str, Any]]:
specs: dict[str, dict[str, Any]] = {}
for alpha_id, function in _PHASE3_FORMULA_FUNCTIONS.items():
meta = ALPHA158_REGISTRY[alpha_id]
call_inputs = _phase3_call_inputs(function)
formula_inputs = _PHASE3_FORMULA_INPUT_OVERRIDES.get(alpha_id, call_inputs)
input_category = _PHASE3_INPUT_CATEGORIES.get(len(call_inputs))
if input_category is None:
raise RuntimeError(f"unsupported formula input count for {alpha_id}")
specs[alpha_id] = {
"name": alpha_id,
"contract_version": ALPHA158_PHASE3_FORMULA_CONTRACT_VERSION,
"formula": meta["formula"],
"category": meta["category"],
"complexity": meta["complexity"],
"parameters": _phase3_string_list(meta, "params", alpha_id),
"description": meta["description"],
"references": _phase3_string_list(meta, "references", alpha_id),
"call_inputs": list(call_inputs),
"formula_inputs": list(formula_inputs),
"input_category": input_category,
}
return specs
def _freeze_phase3_formula_specs(
specs: dict[str, dict[str, Any]],
) -> Mapping[str, Mapping[str, Any]]:
frozen_specs: dict[str, Mapping[str, Any]] = {}
for alpha_id, spec in specs.items():
frozen_specs[alpha_id] = MappingProxyType(
{field: tuple(value) if isinstance(value, list) else value for field, value in spec.items()}
)
return MappingProxyType(frozen_specs)
ALPHA158_PHASE3_FORMULA_SPECS: Mapping[str, Mapping[str, Any]] = (
_freeze_phase3_formula_specs(_build_phase3_formula_specs())
)
def list_phase3_formulas() -> tuple[str, ...]:
"""Return the frozen alpha001-alpha050 formula IDs in stable order."""
return tuple(ALPHA158_PHASE3_FORMULA_SPECS)
def evaluate_phase3_formula(name: str, **inputs: pd.Series) -> pd.Series:
"""Evaluate a Phase 3 formula with an exact, alignment-safe input contract."""
if name not in ALPHA158_PHASE3_FORMULA_SPECS:
raise KeyError(f"formula {name!r} not registered")
spec = ALPHA158_PHASE3_FORMULA_SPECS[name]
required_inputs = cast(tuple[str, ...], spec["call_inputs"])
missing_inputs = [field for field in required_inputs if field not in inputs]
unexpected_inputs = sorted(field for field in inputs if field not in required_inputs)
if missing_inputs or unexpected_inputs:
details: list[str] = []
if missing_inputs:
details.append(f"missing inputs {missing_inputs}")
if unexpected_inputs:
details.append(f"unexpected inputs {unexpected_inputs}")
raise ValueError(f"invalid inputs for {name}: {'; '.join(details)}")
for field in required_inputs:
if not isinstance(inputs[field], pd.Series):
raise TypeError(f"{field} must be a pandas Series")
primary_field = required_inputs[0]
primary = inputs[primary_field]
for field in required_inputs[1:]:
if not primary.index.equals(inputs[field].index):
raise ValueError(f"{field} index must align with {primary_field}")
function = _PHASE3_FORMULA_FUNCTIONS[name]
return function(*(inputs[field] for field in required_inputs))
# ── Phase 4 formula contract: frozen alpha051-alpha100 surface ──────────────
# Phase 4 extends the versioned formula contract without mutating the Phase 3
# catalogue, digest, dispatch surface, or the existing formula functions.
ALPHA158_PHASE4_FORMULA_CONTRACT_VERSION = "1.0.0"
ALPHA158_PHASE4_FORMULA_CATALOG_SHA256 = (
"858daf5e2abab5063fc28fcf7c79936096e2c76dde582fc7bab78b3458f17054"
)
_PHASE4_FORMULA_FUNCTIONS: dict[str, Callable[..., pd.Series]] = {
"alpha_051": alpha_051,
"alpha_052": alpha_052,
"alpha_053": alpha_053,
"alpha_054": alpha_054,
"alpha_055": alpha_055,
"alpha_056": alpha_056,
"alpha_057": alpha_057,
"alpha_058": alpha_058,
"alpha_059": alpha_059,
"alpha_060": alpha_060,
"alpha_061": alpha_061,
"alpha_062": alpha_062,
"alpha_063": alpha_063,
"alpha_064": alpha_064,
"alpha_065": alpha_065,
"alpha_066": alpha_066,
"alpha_067": alpha_067,
"alpha_068": alpha_068,
"alpha_069": alpha_069,
"alpha_070": alpha_070,
"alpha_071": alpha_071,
"alpha_072": alpha_072,
"alpha_073": alpha_073,
"alpha_074": alpha_074,
"alpha_075": alpha_075,
"alpha_076": alpha_076,
"alpha_077": alpha_077,
"alpha_078": alpha_078,
"alpha_079": alpha_079,
"alpha_080": alpha_080,
"alpha_081": alpha_081,
"alpha_082": alpha_082,
"alpha_083": alpha_083,
"alpha_084": alpha_084,
"alpha_085": alpha_085,
"alpha_086": alpha_086,
"alpha_087": alpha_087,
"alpha_088": alpha_088,
"alpha_089": alpha_089,
"alpha_090": alpha_090,
"alpha_091": alpha_091,
"alpha_092": alpha_092,
"alpha_093": alpha_093,
"alpha_094": alpha_094,
"alpha_095": alpha_095,
"alpha_096": alpha_096,
"alpha_097": alpha_097,
"alpha_098": alpha_098,
"alpha_099": alpha_099,
"alpha_100": alpha_100,
}
_PHASE4_INPUT_CATEGORIES = {
1: "single",
2: "pair",
3: "triple",
4: "quadruple",
5: "quintuple",
}
def _build_phase4_formula_specs() -> dict[str, dict[str, Any]]:
specs: dict[str, dict[str, Any]] = {}
for alpha_id, function in _PHASE4_FORMULA_FUNCTIONS.items():
meta = ALPHA158_REGISTRY[alpha_id]
call_inputs = _phase3_call_inputs(function)
formula_inputs = _phase3_string_list(meta, "inputs", alpha_id)
input_category = _PHASE4_INPUT_CATEGORIES.get(len(call_inputs))
if input_category is None:
raise RuntimeError(f"unsupported formula input count for {alpha_id}")
specs[alpha_id] = {
"name": alpha_id,
"contract_version": ALPHA158_PHASE4_FORMULA_CONTRACT_VERSION,
"formula": meta["formula"],
"category": meta["category"],
"complexity": meta["complexity"],
"parameters": _phase3_string_list(meta, "params", alpha_id),
"description": meta["description"],
"references": _phase3_string_list(meta, "references", alpha_id),
"call_inputs": list(call_inputs),
"formula_inputs": formula_inputs,
"input_category": input_category,
}
return specs
ALPHA158_PHASE4_FORMULA_SPECS: Mapping[str, Mapping[str, Any]] = (
_freeze_phase3_formula_specs(_build_phase4_formula_specs())
)
def list_phase4_formulas() -> tuple[str, ...]:
"""Return the frozen alpha051-alpha100 formula IDs in stable order."""
return tuple(ALPHA158_PHASE4_FORMULA_SPECS)
def evaluate_phase4_formula(name: str, **inputs: pd.Series) -> pd.Series:
"""Evaluate a Phase 4 formula with an exact, alignment-safe input contract."""
if name not in ALPHA158_PHASE4_FORMULA_SPECS:
raise KeyError(f"formula {name!r} not registered")
spec = ALPHA158_PHASE4_FORMULA_SPECS[name]
required_inputs = cast(tuple[str, ...], spec["call_inputs"])
missing_inputs = [field for field in required_inputs if field not in inputs]
unexpected_inputs = sorted(field for field in inputs if field not in required_inputs)
if missing_inputs or unexpected_inputs:
details: list[str] = []
if missing_inputs:
details.append(f"missing inputs {missing_inputs}")
if unexpected_inputs:
details.append(f"unexpected inputs {unexpected_inputs}")
raise ValueError(f"invalid inputs for {name}: {'; '.join(details)}")
for field in required_inputs:
if not isinstance(inputs[field], pd.Series):
raise TypeError(f"{field} must be a pandas Series")
primary_field = required_inputs[0]
primary = inputs[primary_field]
for field in required_inputs[1:]:
if not primary.index.equals(inputs[field].index):
raise ValueError(f"{field} index must align with {primary_field}")
function = _PHASE4_FORMULA_FUNCTIONS[name]
return function(*(inputs[field] for field in required_inputs))
# ── Phase 5 formula contract: frozen alpha101-alpha150 surface ──────────────
# Phase 5 extends the versioned formula contract without mutating any earlier
# catalogue, digest, dispatch surface, or existing formula implementation.
ALPHA158_PHASE5_FORMULA_CONTRACT_VERSION = "1.0.0"
ALPHA158_PHASE5_FORMULA_CATALOG_SHA256 = (
"3368796169c9fbd39c4a34ea137e569964b15882fbf4de25124790d548db6533"
)
_PHASE5_FORMULA_FUNCTIONS: dict[str, Callable[..., pd.Series]] = {
"alpha_101": alpha_101,
"alpha_102": alpha_102,
"alpha_103": alpha_103,
"alpha_104": alpha_104,
"alpha_105": alpha_105,
"alpha_106": alpha_106,
"alpha_107": alpha_107,
"alpha_108": alpha_108,
"alpha_109": alpha_109,
"alpha_110": alpha_110,
"alpha_111": alpha_111,
"alpha_112": alpha_112,
"alpha_113": alpha_113,
"alpha_114": alpha_114,
"alpha_115": alpha_115,
"alpha_116": alpha_116,
"alpha_117": alpha_117,
"alpha_118": alpha_118,
"alpha_119": alpha_119,
"alpha_120": alpha_120,
"alpha_121": alpha_121,
"alpha_122": alpha_122,
"alpha_123": alpha_123,
"alpha_124": alpha_124,
"alpha_125": alpha_125,
"alpha_126": alpha_126,
"alpha_127": alpha_127,
"alpha_128": alpha_128,
"alpha_129": alpha_129,
"alpha_130": alpha_130,
"alpha_131": alpha_131,
"alpha_132": alpha_132,
"alpha_133": alpha_133,
"alpha_134": alpha_134,
"alpha_135": alpha_135,
"alpha_136": alpha_136,
"alpha_137": alpha_137,
"alpha_138": alpha_138,
"alpha_139": alpha_139,
"alpha_140": alpha_140,
"alpha_141": alpha_141,
"alpha_142": alpha_142,
"alpha_143": alpha_143,
"alpha_144": alpha_144,
"alpha_145": alpha_145,
"alpha_146": alpha_146,
"alpha_147": alpha_147,
"alpha_148": alpha_148,
"alpha_149": alpha_149,
"alpha_150": alpha_150,
}
def _build_phase5_formula_specs() -> dict[str, dict[str, Any]]:
specs: dict[str, dict[str, Any]] = {}
for alpha_id, function in _PHASE5_FORMULA_FUNCTIONS.items():
meta = ALPHA158_REGISTRY[alpha_id]
call_inputs = _phase3_call_inputs(function)
formula_inputs = _phase3_string_list(meta, "inputs", alpha_id)
input_category = _PHASE4_INPUT_CATEGORIES.get(len(call_inputs))
if input_category is None:
raise RuntimeError(f"unsupported formula input count for {alpha_id}")
specs[alpha_id] = {
"name": alpha_id,
"contract_version": ALPHA158_PHASE5_FORMULA_CONTRACT_VERSION,
"formula": meta["formula"],
"category": meta["category"],
"complexity": meta["complexity"],
"parameters": _phase3_string_list(meta, "params", alpha_id),
"description": meta["description"],
"references": _phase3_string_list(meta, "references", alpha_id),
"call_inputs": list(call_inputs),
"formula_inputs": formula_inputs,
"input_category": input_category,
}
return specs
ALPHA158_PHASE5_FORMULA_SPECS: Mapping[str, Mapping[str, Any]] = (
_freeze_phase3_formula_specs(_build_phase5_formula_specs())
)
def list_phase5_formulas() -> tuple[str, ...]:
"""Return the frozen alpha101-alpha150 formula IDs in stable order."""
return tuple(ALPHA158_PHASE5_FORMULA_SPECS)
def evaluate_phase5_formula(name: str, **inputs: pd.Series) -> pd.Series:
"""Evaluate a Phase 5 formula with an exact, alignment-safe input contract."""
if name not in ALPHA158_PHASE5_FORMULA_SPECS:
raise KeyError(f"formula {name!r} not registered")
spec = ALPHA158_PHASE5_FORMULA_SPECS[name]
required_inputs = cast(tuple[str, ...], spec["call_inputs"])
missing_inputs = [field for field in required_inputs if field not in inputs]
unexpected_inputs = sorted(field for field in inputs if field not in required_inputs)
if missing_inputs or unexpected_inputs:
details: list[str] = []
if missing_inputs:
details.append(f"missing inputs {missing_inputs}")
if unexpected_inputs:
details.append(f"unexpected inputs {unexpected_inputs}")
raise ValueError(f"invalid inputs for {name}: {'; '.join(details)}")
for field in required_inputs:
if not isinstance(inputs[field], pd.Series):
raise TypeError(f"{field} must be a pandas Series")
primary_field = required_inputs[0]
primary = inputs[primary_field]
for field in required_inputs[1:]:
if len(inputs[field]) != len(primary):
raise ValueError(f"{field} length must match {primary_field}")
if not primary.index.equals(inputs[field].index):
raise ValueError(f"{field} index must align with {primary_field}")
function = _PHASE5_FORMULA_FUNCTIONS[name]
return function(*(inputs[field] for field in required_inputs))
# ── Phase 6 formula contract: frozen alpha151-alpha158 surface ──────────────
# Phase 6 completes the versioned formula contract without mutating any
# earlier catalogue, digest, dispatch surface, or formula implementation.
ALPHA158_PHASE6_FORMULA_CONTRACT_VERSION = "1.0.0"
ALPHA158_PHASE6_FORMULA_CATALOG_SHA256 = (
"70ffae16ca6cbb59a1e8644f5cdeec1d291fa75ff8b5647fb534af0083408219"
)
_PHASE6_FORMULA_FUNCTIONS: dict[str, Callable[..., pd.Series]] = {
"alpha_151": alpha_151,
"alpha_152": alpha_152,
"alpha_153": alpha_153,
"alpha_154": alpha_154,
"alpha_155": alpha_155,
"alpha_156": alpha_156,
"alpha_157": alpha_157,
"alpha_158": alpha_158,
}
def _build_phase6_formula_specs() -> dict[str, dict[str, Any]]:
specs: dict[str, dict[str, Any]] = {}
for alpha_id, function in _PHASE6_FORMULA_FUNCTIONS.items():
meta = ALPHA158_REGISTRY[alpha_id]
call_inputs = _phase3_call_inputs(function)
formula_inputs = _phase3_string_list(meta, "inputs", alpha_id)
input_category = _PHASE4_INPUT_CATEGORIES.get(len(call_inputs))
if input_category is None:
raise RuntimeError(f"unsupported formula input count for {alpha_id}")
specs[alpha_id] = {
"name": alpha_id,
"contract_version": ALPHA158_PHASE6_FORMULA_CONTRACT_VERSION,
"formula": meta["formula"],
"category": meta["category"],
"complexity": meta["complexity"],
"parameters": _phase3_string_list(meta, "params", alpha_id),
"description": meta["description"],
"references": _phase3_string_list(meta, "references", alpha_id),
"call_inputs": list(call_inputs),
"formula_inputs": formula_inputs,
"input_category": input_category,
}
return specs
ALPHA158_PHASE6_FORMULA_SPECS: Mapping[str, Mapping[str, Any]] = (
_freeze_phase3_formula_specs(_build_phase6_formula_specs())
)
def list_phase6_formulas() -> tuple[str, ...]:
"""Return the frozen alpha151-alpha158 formula IDs in stable order."""
return tuple(ALPHA158_PHASE6_FORMULA_SPECS)
def evaluate_phase6_formula(name: str, **inputs: pd.Series) -> pd.Series:
"""Evaluate a Phase 6 formula with an exact, alignment-safe input contract."""
if name not in ALPHA158_PHASE6_FORMULA_SPECS:
raise KeyError(f"formula {name!r} not registered")
spec = ALPHA158_PHASE6_FORMULA_SPECS[name]
required_inputs = cast(tuple[str, ...], spec["call_inputs"])
missing_inputs = [field for field in required_inputs if field not in inputs]
unexpected_inputs = sorted(field for field in inputs if field not in required_inputs)
if missing_inputs or unexpected_inputs:
details: list[str] = []
if missing_inputs:
details.append(f"missing inputs {missing_inputs}")
if unexpected_inputs:
details.append(f"unexpected inputs {unexpected_inputs}")
raise ValueError(f"invalid inputs for {name}: {'; '.join(details)}")
for field in required_inputs:
if not isinstance(inputs[field], pd.Series):
raise TypeError(f"{field} must be a pandas Series")
primary_field = required_inputs[0]
primary = inputs[primary_field]
for field in required_inputs[1:]:
if len(inputs[field]) != len(primary):
raise ValueError(f"{field} length must match {primary_field}")
if not primary.index.equals(inputs[field].index):
raise ValueError(f"{field} index must align with {primary_field}")
function = _PHASE6_FORMULA_FUNCTIONS[name]
return function(*(inputs[field] for field in required_inputs))
__all__ = [ __all__ = [
"rank", "rank",
"delta", "delta",
@@ -3653,34 +2755,6 @@ __all__ = [
"max_pair", "max_pair",
"min_pair", "min_pair",
"indneutralize", "indneutralize",
"ALPHA158_PHASE1_MAX_WINDOW",
"ALPHA158_PHASE1_OPERATOR_SPECS",
"list_phase1_operators",
"evaluate_phase1_operator",
"ALPHA158_PHASE2_MAX_WINDOW",
"ALPHA158_PHASE2_OPERATOR_SPECS",
"list_phase2_operators",
"evaluate_phase2_operator",
"ALPHA158_PHASE3_FORMULA_CONTRACT_VERSION",
"ALPHA158_PHASE3_FORMULA_CATALOG_SHA256",
"ALPHA158_PHASE3_FORMULA_SPECS",
"list_phase3_formulas",
"evaluate_phase3_formula",
"ALPHA158_PHASE4_FORMULA_CONTRACT_VERSION",
"ALPHA158_PHASE4_FORMULA_CATALOG_SHA256",
"ALPHA158_PHASE4_FORMULA_SPECS",
"list_phase4_formulas",
"evaluate_phase4_formula",
"ALPHA158_PHASE5_FORMULA_CONTRACT_VERSION",
"ALPHA158_PHASE5_FORMULA_CATALOG_SHA256",
"ALPHA158_PHASE5_FORMULA_SPECS",
"list_phase5_formulas",
"evaluate_phase5_formula",
"ALPHA158_PHASE6_FORMULA_CONTRACT_VERSION",
"ALPHA158_PHASE6_FORMULA_CATALOG_SHA256",
"ALPHA158_PHASE6_FORMULA_SPECS",
"list_phase6_formulas",
"evaluate_phase6_formula",
"alpha_001", "alpha_001",
"alpha_002", "alpha_002",
"alpha_003", "alpha_003",
-540
View File
@@ -1,540 +0,0 @@
"""Versioned, risk-gated, paper-only quantitative research vertical slice.
The module composes existing ``quant_engine`` calculations into a small set of
storage-neutral governance contracts. It never reads a database, calls a data
vendor, loads credentials, or routes an order to a broker.
"""
from __future__ import annotations
import hashlib
import json
import math
import re
from collections.abc import Mapping
from dataclasses import asdict, dataclass
from datetime import UTC, datetime
from enum import StrEnum
from types import MappingProxyType
import pandas as pd
from quant_engine.execution import ExecutionConfig
from quant_engine.research_pipeline import FactorBacktestResult, run_factor_backtest_research
_SHA256 = re.compile(r"^[0-9a-f]{64}$")
_GIT_SHA = re.compile(r"^[0-9a-f]{40}$")
__all__ = [
"DatasetSnapshot",
"FactorVersion",
"StrategyStage",
"StrategyVersion",
"BacktestRun",
"PortfolioTarget",
"RiskPolicy",
"RiskDecisionStatus",
"RiskDecision",
"PaperOrderIntent",
"GovernedFactorSliceResult",
"evaluate_portfolio_risk",
"create_paper_order_intent",
"run_governed_factor_slice",
]
def _required_text(value: str, name: str) -> str:
normalized = value.strip()
if not normalized:
raise ValueError(f"{name} must be non-empty")
return normalized
def _aware_utc(value: datetime, name: str) -> datetime:
if not isinstance(value, datetime) or value.tzinfo is None or value.utcoffset() is None:
raise ValueError(f"{name} must be timezone-aware")
return value.astimezone(UTC)
def _canonical_value(value: object) -> object:
if value is None or isinstance(value, str | bool | int):
return value
if isinstance(value, float):
if math.isnan(value):
return "NaN"
if math.isinf(value):
return "Infinity" if value > 0 else "-Infinity"
return value
if isinstance(value, datetime):
return _aware_utc(value, "datetime").isoformat()
if isinstance(value, StrEnum):
return value.value
if isinstance(value, Mapping):
return {
str(key): _canonical_value(item)
for key, item in sorted(value.items(), key=lambda pair: str(pair[0]))
}
if isinstance(value, list | tuple):
return [_canonical_value(item) for item in value]
raise TypeError(f"unsupported canonical value: {type(value).__name__}")
def _stable_id(prefix: str, payload: Mapping[str, object]) -> str:
encoded = json.dumps(
_canonical_value(payload),
ensure_ascii=False,
sort_keys=True,
separators=(",", ":"),
allow_nan=False,
).encode("utf-8")
return f"{prefix}:{hashlib.sha256(encoded).hexdigest()}"
def _immutable_weights(values: Mapping[str, float]) -> Mapping[str, float]:
normalized: dict[str, float] = {}
for raw_asset, raw_weight in values.items():
asset = _required_text(str(raw_asset), "asset_id")
weight = float(raw_weight)
if not math.isfinite(weight) or weight < 0:
raise ValueError("portfolio weights must be finite and non-negative")
if asset in normalized:
raise ValueError(f"duplicate asset_id: {asset}")
normalized[asset] = weight
return MappingProxyType(dict(sorted(normalized.items())))
@dataclass(frozen=True, slots=True)
class DatasetSnapshot:
"""Point-in-time identity for caller-supplied research data."""
snapshot_id: str
schema_version: str
content_sha256: str
effective_at: datetime
available_at: datetime
ingested_at: datetime
def __post_init__(self) -> None:
object.__setattr__(self, "snapshot_id", _required_text(self.snapshot_id, "snapshot_id"))
object.__setattr__(
self, "schema_version", _required_text(self.schema_version, "schema_version")
)
digest = self.content_sha256.strip().lower()
if not _SHA256.fullmatch(digest):
raise ValueError("content_sha256 must be a lowercase SHA-256 digest")
object.__setattr__(self, "content_sha256", digest)
effective_at = _aware_utc(self.effective_at, "effective_at")
available_at = _aware_utc(self.available_at, "available_at")
ingested_at = _aware_utc(self.ingested_at, "ingested_at")
if not effective_at <= available_at <= ingested_at:
raise ValueError("timestamps must satisfy effective_at <= available_at <= ingested_at")
object.__setattr__(self, "effective_at", effective_at)
object.__setattr__(self, "available_at", available_at)
object.__setattr__(self, "ingested_at", ingested_at)
@dataclass(frozen=True, slots=True)
class FactorVersion:
"""Versioned factor definition identity without factor implementation duplication."""
factor_id: str
version: str
definition_sha256: str
dataset_schema_version: str
def __post_init__(self) -> None:
object.__setattr__(self, "factor_id", _required_text(self.factor_id, "factor_id"))
object.__setattr__(self, "version", _required_text(self.version, "version"))
object.__setattr__(
self,
"dataset_schema_version",
_required_text(self.dataset_schema_version, "dataset_schema_version"),
)
digest = self.definition_sha256.strip().lower()
if not _SHA256.fullmatch(digest):
raise ValueError("definition_sha256 must be a lowercase SHA-256 digest")
object.__setattr__(self, "definition_sha256", digest)
@property
def version_id(self) -> str:
return f"{self.factor_id}@{self.version}"
class StrategyStage(StrEnum):
DRAFT = "Draft"
RESEARCH = "Research"
VALIDATED = "Validated"
APPROVED = "Approved"
PAPER = "Paper"
LIVE = "Live"
PAUSED = "Paused"
RETIRED = "Retired"
@dataclass(frozen=True, slots=True)
class StrategyVersion:
"""Strategy lineage and lifecycle state used by the governed slice."""
strategy_id: str
version: str
factor_version_id: str
stage: StrategyStage
def __post_init__(self) -> None:
object.__setattr__(self, "strategy_id", _required_text(self.strategy_id, "strategy_id"))
object.__setattr__(self, "version", _required_text(self.version, "version"))
object.__setattr__(
self,
"factor_version_id",
_required_text(self.factor_version_id, "factor_version_id"),
)
if not isinstance(self.stage, StrategyStage):
raise TypeError("stage must be a StrategyStage")
@property
def version_id(self) -> str:
return f"{self.strategy_id}@{self.version}"
@dataclass(frozen=True, slots=True)
class BacktestRun:
run_id: str
dataset_snapshot_id: str
factor_version_id: str
strategy_version_id: str
code_revision: str
config_hash: str
created_at: datetime
@dataclass(frozen=True, slots=True)
class PortfolioTarget:
target_id: str
backtest_run_id: str
dataset_snapshot_id: str
weights: Mapping[str, float]
created_at: datetime
def __post_init__(self) -> None:
object.__setattr__(self, "weights", _immutable_weights(self.weights))
@property
def gross_exposure(self) -> float:
return float(sum(abs(weight) for weight in self.weights.values()))
@dataclass(frozen=True, slots=True)
class RiskPolicy:
policy_id: str
max_gross_exposure: float
max_single_asset_weight: float
max_positions: int
def __post_init__(self) -> None:
object.__setattr__(self, "policy_id", _required_text(self.policy_id, "policy_id"))
if not math.isfinite(self.max_gross_exposure) or self.max_gross_exposure <= 0:
raise ValueError("max_gross_exposure must be positive and finite")
if not math.isfinite(self.max_single_asset_weight) or not (
0 < self.max_single_asset_weight <= 1
):
raise ValueError("max_single_asset_weight must be in (0, 1]")
if (
isinstance(self.max_positions, bool)
or not isinstance(self.max_positions, int)
or self.max_positions <= 0
):
raise ValueError("max_positions must be a positive integer")
class RiskDecisionStatus(StrEnum):
APPROVED = "approved"
REJECTED = "rejected"
@dataclass(frozen=True, slots=True)
class RiskDecision:
decision_id: str
portfolio_target_id: str
policy_id: str
status: RiskDecisionStatus
reasons: tuple[str, ...]
created_at: datetime
@dataclass(frozen=True, slots=True, init=False)
class PaperOrderIntent:
intent_id: str
portfolio_target_id: str
risk_decision_id: str
environment: str
target_weights: Mapping[str, float]
created_at: datetime
def __init__(self, target: PortfolioTarget, decision: RiskDecision) -> None:
if decision.status is not RiskDecisionStatus.APPROVED:
raise ValueError("an approved risk decision is required")
if decision.portfolio_target_id != target.target_id:
raise ValueError("risk decision does not bind the supplied portfolio target")
target_weights = _immutable_weights(target.weights)
payload = {
"portfolio_target_id": target.target_id,
"risk_decision_id": decision.decision_id,
"environment": "paper",
"target_weights": target_weights,
"created_at": decision.created_at,
}
object.__setattr__(self, "intent_id", _stable_id("order-intent", payload))
object.__setattr__(self, "portfolio_target_id", target.target_id)
object.__setattr__(self, "risk_decision_id", decision.decision_id)
object.__setattr__(self, "environment", "paper")
object.__setattr__(self, "target_weights", target_weights)
object.__setattr__(self, "created_at", decision.created_at)
@dataclass(frozen=True, slots=True)
class GovernedFactorSliceResult:
dataset_snapshot: DatasetSnapshot
factor_version: FactorVersion
strategy_version: StrategyVersion
research_result: FactorBacktestResult
backtest_run: BacktestRun
portfolio_target: PortfolioTarget
risk_decision: RiskDecision
order_intent: PaperOrderIntent | None
def evaluate_portfolio_risk(
target: PortfolioTarget,
policy: RiskPolicy,
*,
created_at: datetime | None = None,
) -> RiskDecision:
"""Apply fail-closed paper risk limits to a versioned target portfolio."""
decision_time = target.created_at if created_at is None else _aware_utc(created_at, "created_at")
reasons: list[str] = []
if target.gross_exposure > policy.max_gross_exposure + 1e-12:
reasons.append(
f"gross exposure {target.gross_exposure:.12g} exceeds "
f"limit {policy.max_gross_exposure:.12g}"
)
positive_weights = [weight for weight in target.weights.values() if weight > 0]
largest_weight = max(positive_weights, default=0.0)
if largest_weight > policy.max_single_asset_weight + 1e-12:
reasons.append(
f"single-asset weight {largest_weight:.12g} exceeds "
f"limit {policy.max_single_asset_weight:.12g}"
)
if len(positive_weights) > policy.max_positions:
reasons.append(
f"position count {len(positive_weights)} exceeds limit {policy.max_positions}"
)
status = RiskDecisionStatus.REJECTED if reasons else RiskDecisionStatus.APPROVED
payload = {
"portfolio_target_id": target.target_id,
"policy_id": policy.policy_id,
"status": status,
"reasons": reasons,
"created_at": decision_time,
}
return RiskDecision(
decision_id=_stable_id("risk-decision", payload),
portfolio_target_id=target.target_id,
policy_id=policy.policy_id,
status=status,
reasons=tuple(reasons),
created_at=decision_time,
)
def create_paper_order_intent(
target: PortfolioTarget,
decision: RiskDecision,
) -> PaperOrderIntent:
"""Create a paper-only intent after an exact, approved risk decision."""
return PaperOrderIntent(target, decision)
def run_governed_factor_slice(
*,
factor_scores: pd.DataFrame,
execution_prices: pd.DataFrame,
valuation_prices: pd.DataFrame,
dataset_snapshot: DatasetSnapshot,
factor_version: FactorVersion,
strategy_version: StrategyVersion,
risk_policy: RiskPolicy,
code_revision: str,
created_at: datetime,
top_k: int,
execution_price_field: str,
valuation_price_field: str,
lag_sessions: int = 1,
gross_exposure: float = 1.0,
largest: bool = True,
initial_cash: float = 1_000_000.0,
execution_config: ExecutionConfig | None = None,
) -> GovernedFactorSliceResult:
"""Run snapshot → factor → strategy → backtest → target → risk → paper intent."""
if strategy_version.factor_version_id != factor_version.version_id:
raise ValueError("strategy factor lineage does not match factor_version")
if dataset_snapshot.schema_version != factor_version.dataset_schema_version:
raise ValueError("factor dataset schema does not match dataset snapshot")
if strategy_version.stage not in {StrategyStage.APPROVED, StrategyStage.PAPER}:
raise ValueError("strategy must be in Approved or Paper stage")
normalized_revision = code_revision.strip().lower()
if not _GIT_SHA.fullmatch(normalized_revision):
raise ValueError("code_revision must be a lowercase 40-character Git SHA")
run_time = _aware_utc(created_at, "created_at")
if dataset_snapshot.available_at > run_time or dataset_snapshot.ingested_at > run_time:
raise ValueError("dataset snapshot must be available before the research run")
if not factor_scores.empty:
latest_decision = pd.Timestamp(factor_scores.index.max())
latest_decision_date = latest_decision.date()
if latest_decision_date > run_time.date():
raise ValueError("factor_scores contain future decision dates")
config = ExecutionConfig() if execution_config is None else execution_config
if not isinstance(config, ExecutionConfig):
raise TypeError("execution_config must be an ExecutionConfig")
parameters: dict[str, object] = {
"top_k": top_k,
"execution_price_field": execution_price_field,
"valuation_price_field": valuation_price_field,
"lag_sessions": lag_sessions,
"gross_exposure": gross_exposure,
"largest": largest,
"initial_cash": initial_cash,
"execution_config": asdict(config),
}
config_hash = _stable_id("config", parameters).split(":", maxsplit=1)[1]
run_payload = {
"dataset_snapshot_id": dataset_snapshot.snapshot_id,
"dataset_content_sha256": dataset_snapshot.content_sha256,
"factor_version_id": factor_version.version_id,
"factor_definition_sha256": factor_version.definition_sha256,
"strategy_version_id": strategy_version.version_id,
"code_revision": normalized_revision,
"config_hash": config_hash,
"created_at": run_time,
}
backtest_run = BacktestRun(
run_id=_stable_id("backtest-run", run_payload),
dataset_snapshot_id=dataset_snapshot.snapshot_id,
factor_version_id=factor_version.version_id,
strategy_version_id=strategy_version.version_id,
code_revision=normalized_revision,
config_hash=config_hash,
created_at=run_time,
)
research_result = run_factor_backtest_research(
factor_scores,
execution_prices,
valuation_prices,
top_k=top_k,
execution_price_field=execution_price_field,
valuation_price_field=valuation_price_field,
lag_sessions=lag_sessions,
gross_exposure=gross_exposure,
largest=largest,
initial_cash=initial_cash,
config=config,
)
if research_result.schedule.decision_weights.empty:
raise ValueError("governed slice requires at least one target portfolio")
final_weights = {
str(asset): float(weight)
for asset, weight in research_result.schedule.decision_weights.iloc[-1].items()
}
target_payload = {
"backtest_run_id": backtest_run.run_id,
"dataset_snapshot_id": dataset_snapshot.snapshot_id,
"weights": final_weights,
"created_at": run_time,
}
portfolio_target = PortfolioTarget(
target_id=_stable_id("portfolio-target", target_payload),
backtest_run_id=backtest_run.run_id,
dataset_snapshot_id=dataset_snapshot.snapshot_id,
weights=final_weights,
created_at=run_time,
)
risk_decision = evaluate_portfolio_risk(
portfolio_target,
risk_policy,
created_at=run_time,
)
order_intent = (
create_paper_order_intent(portfolio_target, risk_decision)
if risk_decision.status is RiskDecisionStatus.APPROVED
else None
)
return GovernedFactorSliceResult(
dataset_snapshot=dataset_snapshot,
factor_version=factor_version,
strategy_version=strategy_version,
research_result=research_result,
backtest_run=backtest_run,
portfolio_target=portfolio_target,
risk_decision=risk_decision,
order_intent=order_intent,
)
def _demo() -> Mapping[str, object]:
dates = pd.date_range("2026-01-05", periods=4, freq="B")
scores = pd.DataFrame({"A": [2.0, 1.0], "B": [1.0, 2.0]}, index=dates[:2])
opens = pd.DataFrame({"A": [10.0, 10.0, 10.2, 10.3], "B": [20.0, 20.0, 20.5, 21.0]}, index=dates)
result = run_governed_factor_slice(
factor_scores=scores,
execution_prices=opens,
valuation_prices=opens * 1.01,
dataset_snapshot=DatasetSnapshot(
snapshot_id="dataset:architecture-smoke-v1",
schema_version="1.0.0",
content_sha256="a" * 64,
effective_at=datetime(2026, 1, 8, 7, tzinfo=UTC),
available_at=datetime(2026, 1, 8, 8, tzinfo=UTC),
ingested_at=datetime(2026, 1, 8, 8, 5, tzinfo=UTC),
),
factor_version=FactorVersion(
factor_id="factor:architecture-smoke",
version="1.0.0",
definition_sha256="b" * 64,
dataset_schema_version="1.0.0",
),
strategy_version=StrategyVersion(
strategy_id="strategy:architecture-smoke",
version="1.0.0",
factor_version_id="factor:architecture-smoke@1.0.0",
stage=StrategyStage.APPROVED,
),
risk_policy=RiskPolicy(
policy_id="risk:architecture-smoke@1.0.0",
max_gross_exposure=1.0,
max_single_asset_weight=0.6,
max_positions=10,
),
code_revision="c" * 40,
created_at=datetime(2026, 1, 9, 1, tzinfo=UTC),
top_k=2,
execution_price_field="open",
valuation_price_field="close",
execution_config=ExecutionConfig(
commission_bps=0,
stamp_tax_bps=0,
slippage_bps=0,
min_trade_amount=0,
),
)
return {
"backtest_run_id": result.backtest_run.run_id,
"portfolio_target_id": result.portfolio_target.target_id,
"risk_decision": result.risk_decision.status.value,
"order_intent_id": result.order_intent.intent_id if result.order_intent else None,
"environment": result.order_intent.environment if result.order_intent else None,
}
if __name__ == "__main__":
print(json.dumps(_demo(), ensure_ascii=False, sort_keys=True))
-925
View File
@@ -6,23 +6,8 @@ import numpy as np
import pandas as pd import pandas as pd
import pytest import pytest
import quant_engine.alpha_factors as alpha_factors_module
from quant_engine.alpha_factors import ( from quant_engine.alpha_factors import (
ALPHA158_REGISTRY, ALPHA158_REGISTRY,
ALPHA158_PHASE1_OPERATOR_SPECS,
ALPHA158_PHASE2_OPERATOR_SPECS,
ALPHA158_PHASE3_FORMULA_CATALOG_SHA256,
ALPHA158_PHASE3_FORMULA_CONTRACT_VERSION,
ALPHA158_PHASE3_FORMULA_SPECS,
ALPHA158_PHASE4_FORMULA_CATALOG_SHA256,
ALPHA158_PHASE4_FORMULA_CONTRACT_VERSION,
ALPHA158_PHASE4_FORMULA_SPECS,
ALPHA158_PHASE5_FORMULA_CATALOG_SHA256,
ALPHA158_PHASE5_FORMULA_CONTRACT_VERSION,
ALPHA158_PHASE5_FORMULA_SPECS,
ALPHA158_PHASE6_FORMULA_CATALOG_SHA256,
ALPHA158_PHASE6_FORMULA_CONTRACT_VERSION,
ALPHA158_PHASE6_FORMULA_SPECS,
alpha_001, alpha_001,
alpha_002, alpha_002,
alpha_003, alpha_003,
@@ -181,18 +166,6 @@ from quant_engine.alpha_factors import (
alpha_156, alpha_156,
alpha_157, alpha_157,
alpha_158, alpha_158,
evaluate_phase1_operator,
evaluate_phase2_operator,
evaluate_phase3_formula,
evaluate_phase4_formula,
evaluate_phase5_formula,
evaluate_phase6_formula,
list_phase1_operators,
list_phase2_operators,
list_phase3_formulas,
list_phase4_formulas,
list_phase5_formulas,
list_phase6_formulas,
correlation, correlation,
covariance, covariance,
decay_linear, decay_linear,
@@ -1259,901 +1232,3 @@ def test_parse_alpha_formula_round_trip_jsonb():
serialized = json.dumps(parsed) serialized = json.dumps(parsed)
assert isinstance(serialized, str) assert isinstance(serialized, str)
assert "ts_rank" in serialized assert "ts_rank" in serialized
# ── v1.2.0 Phase 1: deterministic operator dispatch contract ──────────────
def test_phase1_operator_catalog_is_explicit_and_serializable():
"""Phase 1 exposes a stable, JSON-friendly catalog for downstream callers."""
import json
expected = {
"rank",
"delta",
"ts_mean",
"ts_std",
"ts_rank",
"correlation",
"ts_min",
"ts_max",
"ts_sum",
"decay_linear",
}
assert set(list_phase1_operators()) == expected
assert set(ALPHA158_PHASE1_OPERATOR_SPECS) == expected
json.dumps(ALPHA158_PHASE1_OPERATOR_SPECS)
for name, spec in ALPHA158_PHASE1_OPERATOR_SPECS.items():
assert spec["name"] == name
assert isinstance(spec["inputs"], list)
assert isinstance(spec["formula"], str)
def test_phase1_unary_operators_preserve_index_and_are_deterministic():
values = pd.Series([1.0, 2.0, 3.0, 4.0], index=["a", "b", "c", "d"])
first = evaluate_phase1_operator("rank", values)
second = evaluate_phase1_operator("rank", values)
pd.testing.assert_series_equal(first, second)
assert first.index.equals(values.index)
assert first.iloc[-1] == pytest.approx(1.0)
@pytest.mark.parametrize(
("name", "window"),
[
("delta", 2),
("ts_mean", 2),
("ts_std", 2),
("ts_rank", 2),
("ts_min", 2),
("ts_max", 2),
("ts_sum", 2),
("decay_linear", 2),
],
)
def test_phase1_windowed_operators_require_explicit_window(name: str, window: int):
values = pd.Series([1.0, 2.0, 3.0, 4.0])
result = evaluate_phase1_operator(name, values, window=window)
assert result.index.equals(values.index)
with pytest.raises(ValueError, match="window"):
evaluate_phase1_operator(name, values)
with pytest.raises(ValueError, match="positive integer"):
evaluate_phase1_operator(name, values, window=1.5) # type: ignore[arg-type]
def test_phase1_binary_correlation_requires_aligned_secondary_input():
values = pd.Series([1.0, 2.0, 3.0, 4.0])
other = pd.Series([4.0, 3.0, 2.0, 1.0])
result = evaluate_phase1_operator("correlation", values, other, window=2)
assert result.iloc[-1] == pytest.approx(-1.0)
with pytest.raises(ValueError, match="secondary"):
evaluate_phase1_operator("correlation", values, window=2)
def test_phase1_dispatch_rejects_unknown_or_unused_arguments():
values = pd.Series([1.0, 2.0, 3.0])
with pytest.raises(KeyError, match="not registered"):
evaluate_phase1_operator("unknown", values)
with pytest.raises(ValueError, match="window"):
evaluate_phase1_operator("rank", values, window=2)
with pytest.raises(ValueError, match="secondary"):
evaluate_phase1_operator("rank", values, values)
def test_phase1_dispatch_rejects_window_above_supported_limit():
values = pd.Series([1.0, 2.0, 3.0])
with pytest.raises(ValueError, match="maximum"):
evaluate_phase1_operator("ts_mean", values, window=2**63)
# ── v1.2.0 Phase 2: cumulative deterministic operator contract ─────────────
def test_phase2_operator_catalog_is_cumulative_stable_and_serializable():
"""Phase 2 exposes all existing building blocks without changing Phase 1."""
import json
phase1 = list_phase1_operators()
expected_phase2 = (
*phase1,
"ts_argmin",
"ts_argmax",
"product",
"returns",
"scale",
"signed_power",
"stddev",
"covariance",
"log",
"abs_series",
"sign",
"max_pair",
"min_pair",
"indneutralize",
)
assert list_phase2_operators() == expected_phase2
assert tuple(ALPHA158_PHASE2_OPERATOR_SPECS) == expected_phase2
assert tuple(ALPHA158_PHASE1_OPERATOR_SPECS) == phase1
json.dumps(ALPHA158_PHASE2_OPERATOR_SPECS)
for name, spec in ALPHA158_PHASE2_OPERATOR_SPECS.items():
assert spec["name"] == name
assert isinstance(spec["inputs"], list)
assert isinstance(spec["parameters"], list)
assert isinstance(spec["formula"], str)
@pytest.mark.parametrize("name", ["ts_argmin", "ts_argmax", "product", "stddev"])
def test_phase2_windowed_unary_dispatch_is_deterministic(name: str):
values = pd.Series([3.0, 1.0, 4.0, 2.0], index=["a", "b", "c", "d"])
first = evaluate_phase2_operator(name, values, window=3)
second = evaluate_phase2_operator(name, values, window=3)
pd.testing.assert_series_equal(first, second)
assert first.index.equals(values.index)
with pytest.raises(ValueError, match="window"):
evaluate_phase2_operator(name, values)
@pytest.mark.parametrize("name", ["returns", "scale", "log", "abs_series", "sign"])
def test_phase2_unary_dispatch_rejects_unused_arguments(name: str):
values = pd.Series([1.0, 2.0, 4.0], index=["a", "b", "c"])
result = evaluate_phase2_operator(name, values)
assert result.index.equals(values.index)
with pytest.raises(ValueError, match="window"):
evaluate_phase2_operator(name, values, window=2)
with pytest.raises(ValueError, match="secondary"):
evaluate_phase2_operator(name, values, secondary=values)
@pytest.mark.parametrize(
("name", "window"),
[("correlation", 2), ("covariance", 2), ("max_pair", None), ("min_pair", None)],
)
def test_phase2_binary_dispatch_requires_aligned_secondary(name: str, window: int | None):
values = pd.Series([1.0, 2.0, 3.0], index=["a", "b", "c"])
secondary = pd.Series([3.0, 2.0, 1.0], index=values.index)
result = evaluate_phase2_operator(name, values, secondary=secondary, window=window)
assert result.index.equals(values.index)
with pytest.raises(ValueError, match="secondary is required"):
evaluate_phase2_operator(name, values, window=window)
with pytest.raises(ValueError, match="secondary index"):
evaluate_phase2_operator(
name,
values,
secondary=secondary.rename(index={"c": "z"}),
window=window,
)
def test_phase2_signed_power_requires_finite_numeric_exponent():
values = pd.Series([-4.0, 0.0, 9.0])
result = evaluate_phase2_operator("signed_power", values, exponent=0.5)
pd.testing.assert_series_equal(result, pd.Series([-2.0, 0.0, 3.0]))
for exponent in (None, True, float("inf"), float("nan"), "2"):
with pytest.raises(ValueError, match="exponent"):
evaluate_phase2_operator( # type: ignore[arg-type]
"signed_power",
values,
exponent=exponent,
)
def test_phase2_indneutralize_requires_aligned_groups():
values = pd.Series([1.0, 3.0, 10.0, 14.0], index=["a", "b", "c", "d"])
groups = pd.Series(["x", "x", "y", "y"], index=values.index)
result = evaluate_phase2_operator("indneutralize", values, groups=groups)
pd.testing.assert_series_equal(result, pd.Series([-1.0, 1.0, -2.0, 2.0], index=values.index))
with pytest.raises(ValueError, match="groups is required"):
evaluate_phase2_operator("indneutralize", values)
with pytest.raises(ValueError, match="groups index"):
evaluate_phase2_operator(
"indneutralize",
values,
groups=groups.rename(index={"d": "z"}),
)
def test_phase2_dispatch_validates_primary_series_and_unused_parameters():
values = pd.Series([1.0, 2.0, 3.0])
with pytest.raises(TypeError, match="series must be a pandas Series"):
evaluate_phase2_operator("rank", [1.0, 2.0, 3.0]) # type: ignore[arg-type]
with pytest.raises(KeyError, match="not registered"):
evaluate_phase2_operator("unknown", values)
with pytest.raises(ValueError, match="exponent"):
evaluate_phase2_operator("rank", values, exponent=2.0)
with pytest.raises(ValueError, match="groups"):
evaluate_phase2_operator("rank", values, groups=pd.Series(["x", "x", "x"]))
with pytest.raises(ValueError, match="maximum"):
evaluate_phase2_operator("product", values, window=253)
# ── Alpha158 Phase 3: versioned alpha001-alpha050 formula contract ──────────
def _phase3_market_inputs() -> dict[str, pd.Series]:
positions = np.arange(80, dtype=float)
index = pd.RangeIndex(len(positions), name="row")
open_ = pd.Series(100.0 + positions * 0.2 + np.sin(positions / 4.0), index=index)
close = pd.Series(100.5 + positions * 0.18 + np.cos(positions / 5.0), index=index)
high = pd.Series(np.maximum(open_, close) + 1.0, index=index)
low = pd.Series(np.minimum(open_, close) - 1.0, index=index)
volume = pd.Series(1_000.0 + positions**1.3 + 20.0 * np.sin(positions / 3.0), index=index)
vwap = (open_ + close + high + low) / 4.0
return {
"open": open_,
"close": close,
"high": high,
"low": low,
"volume": volume,
"vwap": vwap,
}
def test_phase3_formula_catalog_is_versioned_exact_and_content_addressed():
import hashlib
import json
from collections import Counter
expected_ids = tuple(f"alpha_{number:03d}" for number in range(1, 51))
expected_fields = {
"name",
"contract_version",
"formula",
"category",
"complexity",
"parameters",
"description",
"references",
"call_inputs",
"formula_inputs",
"input_category",
}
assert ALPHA158_PHASE3_FORMULA_CONTRACT_VERSION == "1.0.0"
assert list_phase3_formulas() == expected_ids
assert tuple(ALPHA158_PHASE3_FORMULA_SPECS) == expected_ids
assert Counter(
spec["input_category"] for spec in ALPHA158_PHASE3_FORMULA_SPECS.values()
) == {"single": 19, "pair": 26, "triple": 4, "quadruple": 1}
for alpha_id, spec in ALPHA158_PHASE3_FORMULA_SPECS.items():
assert set(spec) == expected_fields
assert spec["name"] == alpha_id
assert spec["contract_version"] == ALPHA158_PHASE3_FORMULA_CONTRACT_VERSION
assert spec["formula"] == ALPHA158_REGISTRY[alpha_id]["formula"]
serializable_specs = {
alpha_id: {
field: list(value) if isinstance(value, tuple) else value
for field, value in spec.items()
}
for alpha_id, spec in ALPHA158_PHASE3_FORMULA_SPECS.items()
}
encoded = json.dumps(
serializable_specs,
sort_keys=True,
separators=(",", ":"),
ensure_ascii=False,
).encode()
assert hashlib.sha256(encoded).hexdigest() == ALPHA158_PHASE3_FORMULA_CATALOG_SHA256
assert ALPHA158_PHASE3_FORMULA_CATALOG_SHA256 == (
"9a3360e5ee77a85a35d3c2fdab1efaa531fa0c263a2cb2bb5b965a1fb96fe1bd"
)
def test_phase3_formula_catalog_is_recursively_immutable():
import operator
with pytest.raises(TypeError):
operator.setitem(ALPHA158_PHASE3_FORMULA_SPECS, "alpha_001", {})
with pytest.raises(TypeError):
operator.setitem(
ALPHA158_PHASE3_FORMULA_SPECS["alpha_001"],
"formula",
"changed",
)
with pytest.raises(TypeError):
operator.setitem(
ALPHA158_PHASE3_FORMULA_SPECS["alpha_001"]["call_inputs"],
0,
"volume",
)
def test_phase3_catalog_freezes_callable_signatures_without_rewriting_formulas():
import inspect
legacy_formula_input_differences = {
"alpha_011": ("close", "high", "low"),
"alpha_035": ("volume",),
"alpha_036": ("close",),
"alpha_040": ("high", "low"),
"alpha_042": ("close",),
"alpha_043": ("volume",),
}
for alpha_id, spec in ALPHA158_PHASE3_FORMULA_SPECS.items():
function = getattr(alpha_factors_module, alpha_id)
signature_inputs = tuple(
"open" if name == "open_" else name
for name in inspect.signature(function).parameters
)
assert spec["call_inputs"] == signature_inputs
assert spec["formula_inputs"] == legacy_formula_input_differences.get(
alpha_id,
signature_inputs,
)
assert ALPHA158_PHASE3_FORMULA_SPECS["alpha_035"]["call_inputs"] == (
"close",
"volume",
)
assert ALPHA158_REGISTRY["alpha_035"]["inputs"] == ["volume"]
def test_phase3_dispatch_matches_all_existing_alpha001_alpha050_functions():
inputs = _phase3_market_inputs()
for alpha_id, spec in ALPHA158_PHASE3_FORMULA_SPECS.items():
call_inputs = spec["call_inputs"]
function = getattr(alpha_factors_module, alpha_id)
expected = function(*(inputs[name] for name in call_inputs))
actual = evaluate_phase3_formula(
alpha_id,
**{name: inputs[name] for name in reversed(call_inputs)},
)
pd.testing.assert_series_equal(actual, expected)
def test_phase3_dispatch_rejects_unknown_missing_extra_and_non_series_inputs():
inputs = _phase3_market_inputs()
with pytest.raises(KeyError, match="not registered"):
evaluate_phase3_formula("alpha_051", close=inputs["close"])
with pytest.raises(ValueError, match=r"missing inputs.*volume"):
evaluate_phase3_formula("alpha_005", close=inputs["close"])
with pytest.raises(ValueError, match=r"unexpected inputs.*vwap"):
evaluate_phase3_formula(
"alpha_005",
close=inputs["close"],
volume=inputs["volume"],
vwap=inputs["vwap"],
)
with pytest.raises(TypeError, match="close must be a pandas Series"):
evaluate_phase3_formula( # type: ignore[arg-type]
"alpha_005",
close=[1.0, 2.0],
volume=inputs["volume"],
)
def test_phase3_dispatch_rejects_implicit_series_alignment():
inputs = _phase3_market_inputs()
misaligned_volume = inputs["volume"].rename(index={79: 80})
with pytest.raises(ValueError, match="volume index must align with close"):
evaluate_phase3_formula(
"alpha_005",
close=inputs["close"],
volume=misaligned_volume,
)
# ── Alpha158 Phase 4: versioned alpha051-alpha100 formula contract ──────────
def test_phase4_formula_catalog_is_versioned_exact_and_content_addressed():
import hashlib
import json
from collections import Counter
expected_ids = tuple(f"alpha_{number:03d}" for number in range(51, 101))
expected_fields = {
"name",
"contract_version",
"formula",
"category",
"complexity",
"parameters",
"description",
"references",
"call_inputs",
"formula_inputs",
"input_category",
}
assert ALPHA158_PHASE4_FORMULA_CONTRACT_VERSION == "1.0.0"
assert list_phase4_formulas() == expected_ids
assert tuple(ALPHA158_PHASE4_FORMULA_SPECS) == expected_ids
assert Counter(
spec["input_category"] for spec in ALPHA158_PHASE4_FORMULA_SPECS.values()
) == {"single": 1, "pair": 27, "triple": 11, "quadruple": 10, "quintuple": 1}
for alpha_id, spec in ALPHA158_PHASE4_FORMULA_SPECS.items():
assert set(spec) == expected_fields
assert spec["name"] == alpha_id
assert spec["contract_version"] == ALPHA158_PHASE4_FORMULA_CONTRACT_VERSION
assert spec["formula"] == ALPHA158_REGISTRY[alpha_id]["formula"]
serializable_specs = {
alpha_id: {
field: list(value) if isinstance(value, tuple) else value
for field, value in spec.items()
}
for alpha_id, spec in ALPHA158_PHASE4_FORMULA_SPECS.items()
}
encoded = json.dumps(
serializable_specs,
sort_keys=True,
separators=(",", ":"),
ensure_ascii=False,
).encode()
assert hashlib.sha256(encoded).hexdigest() == ALPHA158_PHASE4_FORMULA_CATALOG_SHA256
assert ALPHA158_PHASE4_FORMULA_CATALOG_SHA256 == (
"858daf5e2abab5063fc28fcf7c79936096e2c76dde582fc7bab78b3458f17054"
)
def test_phase4_formula_catalog_is_recursively_immutable():
import operator
with pytest.raises(TypeError):
operator.setitem(ALPHA158_PHASE4_FORMULA_SPECS, "alpha_051", {})
with pytest.raises(TypeError):
operator.setitem(
ALPHA158_PHASE4_FORMULA_SPECS["alpha_051"],
"formula",
"changed",
)
with pytest.raises(TypeError):
operator.setitem(
ALPHA158_PHASE4_FORMULA_SPECS["alpha_051"]["call_inputs"],
0,
"volume",
)
def test_phase4_catalog_freezes_callable_and_formula_inputs():
import inspect
for alpha_id, spec in ALPHA158_PHASE4_FORMULA_SPECS.items():
function = getattr(alpha_factors_module, alpha_id)
signature_inputs = tuple(
"open" if name == "open_" else name
for name in inspect.signature(function).parameters
)
assert spec["call_inputs"] == signature_inputs
assert spec["formula_inputs"] == tuple(ALPHA158_REGISTRY[alpha_id]["inputs"])
assert ALPHA158_PHASE4_FORMULA_SPECS["alpha_055"]["call_inputs"] == (
"open",
"high",
"low",
"volume",
"close",
)
def test_phase4_dispatch_matches_all_existing_alpha051_alpha100_functions():
inputs = _phase3_market_inputs()
for alpha_id, spec in ALPHA158_PHASE4_FORMULA_SPECS.items():
call_inputs = spec["call_inputs"]
function = getattr(alpha_factors_module, alpha_id)
expected = function(*(inputs[name] for name in call_inputs))
actual = evaluate_phase4_formula(
alpha_id,
**{name: inputs[name] for name in reversed(call_inputs)},
)
pd.testing.assert_series_equal(actual, expected)
def test_phase4_dispatch_rejects_unknown_missing_extra_and_non_series_inputs():
inputs = _phase3_market_inputs()
with pytest.raises(KeyError, match="not registered"):
evaluate_phase4_formula("alpha_050", close=inputs["close"])
with pytest.raises(ValueError, match=r"missing inputs.*low"):
evaluate_phase4_formula("alpha_051", high=inputs["high"])
with pytest.raises(ValueError, match=r"unexpected inputs.*vwap"):
evaluate_phase4_formula(
"alpha_051",
high=inputs["high"],
low=inputs["low"],
vwap=inputs["vwap"],
)
with pytest.raises(TypeError, match="high must be a pandas Series"):
evaluate_phase4_formula( # type: ignore[arg-type]
"alpha_051",
high=[1.0, 2.0],
low=inputs["low"],
)
def test_phase4_dispatch_rejects_implicit_series_alignment():
inputs = _phase3_market_inputs()
misaligned_low = inputs["low"].rename(index={79: 80})
with pytest.raises(ValueError, match="low index must align with high"):
evaluate_phase4_formula(
"alpha_051",
high=inputs["high"],
low=misaligned_low,
)
# ── Alpha158 Phase 5: versioned alpha101-alpha150 formula contract ──────────
def test_phase5_formula_catalog_is_versioned_exact_and_content_addressed():
import hashlib
import json
from collections import Counter
expected_ids = tuple(f"alpha_{number:03d}" for number in range(101, 151))
expected_fields = {
"name",
"contract_version",
"formula",
"category",
"complexity",
"parameters",
"description",
"references",
"call_inputs",
"formula_inputs",
"input_category",
}
assert ALPHA158_PHASE5_FORMULA_CONTRACT_VERSION == "1.0.0"
assert list_phase5_formulas() == expected_ids
assert tuple(ALPHA158_PHASE5_FORMULA_SPECS) == expected_ids
assert Counter(
spec["input_category"] for spec in ALPHA158_PHASE5_FORMULA_SPECS.values()
) == {"pair": 33, "triple": 14, "quadruple": 3}
for alpha_id, spec in ALPHA158_PHASE5_FORMULA_SPECS.items():
assert set(spec) == expected_fields
assert spec["name"] == alpha_id
assert spec["contract_version"] == ALPHA158_PHASE5_FORMULA_CONTRACT_VERSION
assert spec["formula"] == ALPHA158_REGISTRY[alpha_id]["formula"]
serializable_specs = {
alpha_id: {
field: list(value) if isinstance(value, tuple) else value
for field, value in spec.items()
}
for alpha_id, spec in ALPHA158_PHASE5_FORMULA_SPECS.items()
}
encoded = json.dumps(
serializable_specs,
sort_keys=True,
separators=(",", ":"),
ensure_ascii=False,
).encode()
assert hashlib.sha256(encoded).hexdigest() == ALPHA158_PHASE5_FORMULA_CATALOG_SHA256
assert ALPHA158_PHASE5_FORMULA_CATALOG_SHA256 == (
"3368796169c9fbd39c4a34ea137e569964b15882fbf4de25124790d548db6533"
)
def test_phase5_formula_catalog_is_recursively_immutable():
import operator
with pytest.raises(TypeError):
operator.setitem(ALPHA158_PHASE5_FORMULA_SPECS, "alpha_101", {})
with pytest.raises(TypeError):
operator.setitem(
ALPHA158_PHASE5_FORMULA_SPECS["alpha_101"],
"formula",
"changed",
)
with pytest.raises(TypeError):
operator.setitem(
ALPHA158_PHASE5_FORMULA_SPECS["alpha_101"]["call_inputs"],
0,
"volume",
)
def test_phase5_catalog_freezes_callable_and_formula_inputs():
import inspect
for alpha_id, spec in ALPHA158_PHASE5_FORMULA_SPECS.items():
function = getattr(alpha_factors_module, alpha_id)
signature_inputs = tuple(
"open" if name == "open_" else name
for name in inspect.signature(function).parameters
)
assert spec["call_inputs"] == signature_inputs
assert spec["formula_inputs"] == tuple(ALPHA158_REGISTRY[alpha_id]["inputs"])
assert ALPHA158_PHASE5_FORMULA_SPECS["alpha_101"]["call_inputs"] == (
"close",
"high",
"low",
)
def test_phase5_dispatch_matches_all_existing_alpha101_alpha150_functions():
inputs = _phase3_market_inputs()
for alpha_id, spec in ALPHA158_PHASE5_FORMULA_SPECS.items():
call_inputs = spec["call_inputs"]
function = getattr(alpha_factors_module, alpha_id)
expected = function(*(inputs[name] for name in call_inputs))
actual = evaluate_phase5_formula(
alpha_id,
**{name: inputs[name] for name in reversed(call_inputs)},
)
pd.testing.assert_series_equal(actual, expected)
def test_phase5_dispatch_rejects_unknown_missing_extra_and_non_series_inputs():
inputs = _phase3_market_inputs()
with pytest.raises(KeyError, match="not registered"):
evaluate_phase5_formula("alpha_100", close=inputs["close"])
with pytest.raises(KeyError, match="not registered"):
evaluate_phase5_formula("alpha_151", close=inputs["close"])
with pytest.raises(ValueError, match=r"missing inputs.*low"):
evaluate_phase5_formula(
"alpha_101",
close=inputs["close"],
high=inputs["high"],
)
with pytest.raises(ValueError, match=r"unexpected inputs.*vwap"):
evaluate_phase5_formula(
"alpha_101",
close=inputs["close"],
high=inputs["high"],
low=inputs["low"],
vwap=inputs["vwap"],
)
with pytest.raises(TypeError, match="high must be a pandas Series"):
evaluate_phase5_formula( # type: ignore[arg-type]
"alpha_101",
close=inputs["close"],
high=[1.0, 2.0],
low=inputs["low"],
)
def test_phase5_dispatch_rejects_length_and_index_alignment_errors():
inputs = _phase3_market_inputs()
shorter_low = inputs["low"].iloc[:-1]
misaligned_high = inputs["high"].rename(index={79: 80})
with pytest.raises(ValueError, match="low length must match close"):
evaluate_phase5_formula(
"alpha_101",
close=inputs["close"],
high=inputs["high"],
low=shorter_low,
)
with pytest.raises(ValueError, match="high index must align with close"):
evaluate_phase5_formula(
"alpha_101",
close=inputs["close"],
high=misaligned_high,
low=inputs["low"],
)
def test_phase5_contract_is_publicly_exported():
assert {
"ALPHA158_PHASE5_FORMULA_CONTRACT_VERSION",
"ALPHA158_PHASE5_FORMULA_CATALOG_SHA256",
"ALPHA158_PHASE5_FORMULA_SPECS",
"list_phase5_formulas",
"evaluate_phase5_formula",
} <= set(alpha_factors_module.__all__)
# ── Alpha158 Phase 6: versioned alpha151-alpha158 formula contract ──────────
def test_phase6_formula_catalog_is_versioned_exact_and_content_addressed():
import hashlib
import json
from collections import Counter
expected_ids = tuple(f"alpha_{number:03d}" for number in range(151, 159))
expected_fields = {
"name",
"contract_version",
"formula",
"category",
"complexity",
"parameters",
"description",
"references",
"call_inputs",
"formula_inputs",
"input_category",
}
assert ALPHA158_PHASE6_FORMULA_CONTRACT_VERSION == "1.0.0"
assert list_phase6_formulas() == expected_ids
assert tuple(ALPHA158_PHASE6_FORMULA_SPECS) == expected_ids
assert Counter(
spec["input_category"] for spec in ALPHA158_PHASE6_FORMULA_SPECS.values()
) == {"pair": 6, "triple": 2}
for alpha_id, spec in ALPHA158_PHASE6_FORMULA_SPECS.items():
assert set(spec) == expected_fields
assert spec["name"] == alpha_id
assert spec["contract_version"] == ALPHA158_PHASE6_FORMULA_CONTRACT_VERSION
assert spec["formula"] == ALPHA158_REGISTRY[alpha_id]["formula"]
serializable_specs = {
alpha_id: {
field: list(value) if isinstance(value, tuple) else value
for field, value in spec.items()
}
for alpha_id, spec in ALPHA158_PHASE6_FORMULA_SPECS.items()
}
encoded = json.dumps(
serializable_specs,
sort_keys=True,
separators=(",", ":"),
ensure_ascii=False,
).encode()
assert hashlib.sha256(encoded).hexdigest() == ALPHA158_PHASE6_FORMULA_CATALOG_SHA256
assert ALPHA158_PHASE6_FORMULA_CATALOG_SHA256 == (
"70ffae16ca6cbb59a1e8644f5cdeec1d291fa75ff8b5647fb534af0083408219"
)
def test_phase6_formula_catalog_is_recursively_immutable():
import operator
with pytest.raises(TypeError):
operator.setitem(ALPHA158_PHASE6_FORMULA_SPECS, "alpha_151", {})
with pytest.raises(TypeError):
operator.setitem(
ALPHA158_PHASE6_FORMULA_SPECS["alpha_151"],
"formula",
"changed",
)
with pytest.raises(TypeError):
operator.setitem(
ALPHA158_PHASE6_FORMULA_SPECS["alpha_151"]["call_inputs"],
0,
"volume",
)
def test_phase6_catalog_freezes_callable_and_formula_inputs():
import inspect
for alpha_id, spec in ALPHA158_PHASE6_FORMULA_SPECS.items():
function = getattr(alpha_factors_module, alpha_id)
signature_inputs = tuple(
"open" if name == "open_" else name
for name in inspect.signature(function).parameters
)
assert spec["call_inputs"] == signature_inputs
assert spec["formula_inputs"] == tuple(ALPHA158_REGISTRY[alpha_id]["inputs"])
assert ALPHA158_PHASE6_FORMULA_SPECS["alpha_158"]["call_inputs"] == (
"high",
"low",
"volume",
)
def test_phase6_dispatch_matches_all_existing_alpha151_alpha158_functions():
inputs = _phase3_market_inputs()
for alpha_id, spec in ALPHA158_PHASE6_FORMULA_SPECS.items():
call_inputs = spec["call_inputs"]
function = getattr(alpha_factors_module, alpha_id)
expected = function(*(inputs[name] for name in call_inputs))
actual = evaluate_phase6_formula(
alpha_id,
**{name: inputs[name] for name in reversed(call_inputs)},
)
pd.testing.assert_series_equal(actual, expected)
def test_phase6_dispatch_rejects_unknown_missing_extra_and_non_series_inputs():
inputs = _phase3_market_inputs()
with pytest.raises(KeyError, match="not registered"):
evaluate_phase6_formula("alpha_150", close=inputs["close"])
with pytest.raises(KeyError, match="not registered"):
evaluate_phase6_formula("alpha_159", close=inputs["close"])
with pytest.raises(ValueError, match=r"missing inputs.*volume"):
evaluate_phase6_formula("alpha_151", close=inputs["close"])
with pytest.raises(ValueError, match=r"unexpected inputs.*vwap"):
evaluate_phase6_formula(
"alpha_151",
close=inputs["close"],
volume=inputs["volume"],
vwap=inputs["vwap"],
)
with pytest.raises(TypeError, match="volume must be a pandas Series"):
evaluate_phase6_formula( # type: ignore[arg-type]
"alpha_151",
close=inputs["close"],
volume=[1.0, 2.0],
)
def test_phase6_dispatch_rejects_length_and_index_alignment_errors():
inputs = _phase3_market_inputs()
shorter_volume = inputs["volume"].iloc[:-1]
misaligned_low = inputs["low"].rename(index={79: 80})
with pytest.raises(ValueError, match="volume length must match close"):
evaluate_phase6_formula(
"alpha_151",
close=inputs["close"],
volume=shorter_volume,
)
with pytest.raises(ValueError, match="low index must align with high"):
evaluate_phase6_formula(
"alpha_158",
high=inputs["high"],
low=misaligned_low,
volume=inputs["volume"],
)
def test_phase6_preserves_complete_alpha001_alpha158_formula_body_fingerprint():
import ast
import hashlib
import inspect
import json
import textwrap
fingerprints = {}
for number in range(1, 159):
alpha_id = f"alpha_{number:03d}"
source = textwrap.dedent(inspect.getsource(getattr(alpha_factors_module, alpha_id)))
node = ast.parse(source).body[0]
assert isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef))
body = ast.dump(
ast.Module(body=node.body, type_ignores=[]),
include_attributes=False,
)
fingerprints[alpha_id] = hashlib.sha256(body.encode()).hexdigest()
encoded = json.dumps(
fingerprints,
sort_keys=True,
separators=(",", ":"),
).encode()
assert hashlib.sha256(encoded).hexdigest() == (
"f3ae807983cf8ca083d0b924c3807ffd84a62a5bd354f9efb1a1a16ca40c7da6"
)
def test_phase6_contract_is_publicly_exported():
assert {
"ALPHA158_PHASE6_FORMULA_CONTRACT_VERSION",
"ALPHA158_PHASE6_FORMULA_CATALOG_SHA256",
"ALPHA158_PHASE6_FORMULA_SPECS",
"list_phase6_formulas",
"evaluate_phase6_formula",
} <= set(alpha_factors_module.__all__)
-320
View File
@@ -1,320 +0,0 @@
"""Governed Personal Quant OS vertical-slice contracts."""
from __future__ import annotations
from datetime import UTC, datetime
import pandas as pd
import pytest
from quant_engine.execution import ExecutionConfig
from quant_engine.governed_pipeline import (
DatasetSnapshot,
FactorVersion,
PaperOrderIntent,
RiskDecisionStatus,
RiskPolicy,
StrategyStage,
StrategyVersion,
create_paper_order_intent,
run_governed_factor_slice,
)
def _calendar() -> pd.DatetimeIndex:
return pd.date_range("2026-01-05", periods=4, freq="B")
def _scores() -> pd.DataFrame:
dates = _calendar()
return pd.DataFrame(
{"A": [3.0, 1.0], "B": [2.0, 3.0], "C": [1.0, 2.0]},
index=dates[:2],
)
def _prices() -> tuple[pd.DataFrame, pd.DataFrame]:
dates = _calendar()
opens = pd.DataFrame(
{"A": [10.0, 10.0, 10.2, 10.4], "B": [20.0, 20.0, 20.5, 21.0], "C": [30.0, 30.0, 30.0, 30.0]},
index=dates,
)
closes = opens * 1.01
return opens, closes
def _snapshot() -> DatasetSnapshot:
return DatasetSnapshot(
snapshot_id="dataset:cn-a-daily-20260108-v1",
schema_version="1.0.0",
content_sha256="a" * 64,
effective_at=datetime(2026, 1, 8, 7, tzinfo=UTC),
available_at=datetime(2026, 1, 8, 8, tzinfo=UTC),
ingested_at=datetime(2026, 1, 8, 8, 5, tzinfo=UTC),
)
def _factor() -> FactorVersion:
return FactorVersion(
factor_id="factor:demo-momentum",
version="1.0.0",
definition_sha256="b" * 64,
dataset_schema_version="1.0.0",
)
def _strategy() -> StrategyVersion:
return StrategyVersion(
strategy_id="strategy:demo-top2",
version="1.0.0",
factor_version_id="factor:demo-momentum@1.0.0",
stage=StrategyStage.APPROVED,
)
def _execution_config() -> ExecutionConfig:
return ExecutionConfig(
commission_bps=0,
stamp_tax_bps=0,
slippage_bps=0,
min_trade_amount=0,
)
def test_governed_slice_is_reproducible_and_creates_only_paper_intent() -> None:
opens, closes = _prices()
created_at = datetime(2026, 1, 9, 1, tzinfo=UTC)
policy = RiskPolicy(
policy_id="risk:paper-default@1.0.0",
max_gross_exposure=1.0,
max_single_asset_weight=0.6,
max_positions=10,
)
result = run_governed_factor_slice(
factor_scores=_scores(),
execution_prices=opens,
valuation_prices=closes,
dataset_snapshot=_snapshot(),
factor_version=_factor(),
strategy_version=_strategy(),
risk_policy=policy,
code_revision="c" * 40,
created_at=created_at,
top_k=2,
execution_price_field="open",
valuation_price_field="close",
execution_config=_execution_config(),
)
assert result.backtest_run.dataset_snapshot_id == _snapshot().snapshot_id
assert result.backtest_run.factor_version_id == _factor().version_id
assert result.backtest_run.strategy_version_id == _strategy().version_id
assert result.backtest_run.code_revision == "c" * 40
assert len(result.backtest_run.config_hash) == 64
assert result.portfolio_target.backtest_run_id == result.backtest_run.run_id
assert result.risk_decision.status is RiskDecisionStatus.APPROVED
assert result.risk_decision.portfolio_target_id == result.portfolio_target.target_id
assert result.order_intent is not None
assert result.order_intent.environment == "paper"
assert result.order_intent.risk_decision_id == result.risk_decision.decision_id
assert result.order_intent.portfolio_target_id == result.portfolio_target.target_id
repeated = run_governed_factor_slice(
factor_scores=_scores(),
execution_prices=opens,
valuation_prices=closes,
dataset_snapshot=_snapshot(),
factor_version=_factor(),
strategy_version=_strategy(),
risk_policy=policy,
code_revision="c" * 40,
created_at=created_at,
top_k=2,
execution_price_field="open",
valuation_price_field="close",
execution_config=_execution_config(),
)
assert repeated.backtest_run.run_id == result.backtest_run.run_id
assert repeated.portfolio_target.target_id == result.portfolio_target.target_id
assert repeated.risk_decision.decision_id == result.risk_decision.decision_id
assert repeated.order_intent == result.order_intent
def test_risk_rejection_blocks_order_intent() -> None:
opens, closes = _prices()
result = run_governed_factor_slice(
factor_scores=_scores(),
execution_prices=opens,
valuation_prices=closes,
dataset_snapshot=_snapshot(),
factor_version=_factor(),
strategy_version=_strategy(),
risk_policy=RiskPolicy(
policy_id="risk:no-concentration@1.0.0",
max_gross_exposure=1.0,
max_single_asset_weight=0.4,
max_positions=10,
),
code_revision="c" * 40,
created_at=datetime(2026, 1, 9, 1, tzinfo=UTC),
top_k=2,
execution_price_field="open",
valuation_price_field="close",
execution_config=_execution_config(),
)
assert result.risk_decision.status is RiskDecisionStatus.REJECTED
assert any("single-asset weight" in reason for reason in result.risk_decision.reasons)
assert result.order_intent is None
with pytest.raises(ValueError, match="approved risk decision"):
create_paper_order_intent(result.portfolio_target, result.risk_decision)
with pytest.raises(ValueError, match="approved risk decision"):
PaperOrderIntent(result.portfolio_target, result.risk_decision)
def test_dataset_snapshot_requires_point_in_time_ordering_and_aware_times() -> None:
with pytest.raises(ValueError, match="timezone-aware"):
DatasetSnapshot(
snapshot_id="dataset:invalid",
schema_version="1.0.0",
content_sha256="a" * 64,
effective_at=datetime(2026, 1, 8, 7),
available_at=datetime(2026, 1, 8, 8, tzinfo=UTC),
ingested_at=datetime(2026, 1, 8, 9, tzinfo=UTC),
)
with pytest.raises(ValueError, match="effective_at <= available_at <= ingested_at"):
DatasetSnapshot(
snapshot_id="dataset:invalid",
schema_version="1.0.0",
content_sha256="a" * 64,
effective_at=datetime(2026, 1, 8, 9, tzinfo=UTC),
available_at=datetime(2026, 1, 8, 8, tzinfo=UTC),
ingested_at=datetime(2026, 1, 8, 10, tzinfo=UTC),
)
def test_strategy_factor_lineage_must_match() -> None:
opens, closes = _prices()
mismatched = StrategyVersion(
strategy_id="strategy:demo-top2",
version="1.0.0",
factor_version_id="factor:other@1.0.0",
stage=StrategyStage.APPROVED,
)
with pytest.raises(ValueError, match="factor lineage"):
run_governed_factor_slice(
factor_scores=_scores(),
execution_prices=opens,
valuation_prices=closes,
dataset_snapshot=_snapshot(),
factor_version=_factor(),
strategy_version=mismatched,
risk_policy=RiskPolicy(
policy_id="risk:paper-default@1.0.0",
max_gross_exposure=1.0,
max_single_asset_weight=0.6,
max_positions=10,
),
code_revision="c" * 40,
created_at=datetime(2026, 1, 9, 1, tzinfo=UTC),
top_k=2,
execution_price_field="open",
valuation_price_field="close",
execution_config=_execution_config(),
)
def test_governed_slice_requires_matching_schema_and_snapshot_available_by_run_time() -> None:
opens, closes = _prices()
common = {
"factor_scores": _scores(),
"execution_prices": opens,
"valuation_prices": closes,
"strategy_version": _strategy(),
"risk_policy": RiskPolicy(
policy_id="risk:paper-default@1.0.0",
max_gross_exposure=1.0,
max_single_asset_weight=0.6,
max_positions=10,
),
"code_revision": "c" * 40,
"top_k": 2,
"execution_price_field": "open",
"valuation_price_field": "close",
"execution_config": _execution_config(),
}
with pytest.raises(ValueError, match="dataset schema"):
run_governed_factor_slice(
dataset_snapshot=_snapshot(),
factor_version=FactorVersion(
factor_id="factor:demo-momentum",
version="1.0.0",
definition_sha256="b" * 64,
dataset_schema_version="2.0.0",
),
created_at=datetime(2026, 1, 9, 1, tzinfo=UTC),
**common,
)
with pytest.raises(ValueError, match="available before the research run"):
run_governed_factor_slice(
dataset_snapshot=_snapshot(),
factor_version=_factor(),
created_at=datetime(2026, 1, 8, 7, 30, tzinfo=UTC),
**common,
)
future_scores = _scores()
future_scores.index = pd.date_range("2026-01-12", periods=2, freq="B")
with pytest.raises(ValueError, match="future decision dates"):
run_governed_factor_slice(
dataset_snapshot=_snapshot(),
factor_version=_factor(),
factor_scores=future_scores,
execution_prices=opens,
valuation_prices=closes,
strategy_version=common["strategy_version"],
risk_policy=common["risk_policy"],
code_revision="c" * 40,
created_at=datetime(2026, 1, 9, 1, tzinfo=UTC),
top_k=2,
execution_price_field="open",
valuation_price_field="close",
execution_config=_execution_config(),
)
def test_paper_intent_requires_approved_strategy_stage() -> None:
opens, closes = _prices()
validated = StrategyVersion(
strategy_id="strategy:demo-top2",
version="1.0.0",
factor_version_id=_factor().version_id,
stage=StrategyStage.VALIDATED,
)
with pytest.raises(ValueError, match="Approved or Paper"):
run_governed_factor_slice(
factor_scores=_scores(),
execution_prices=opens,
valuation_prices=closes,
dataset_snapshot=_snapshot(),
factor_version=_factor(),
strategy_version=validated,
risk_policy=RiskPolicy(
policy_id="risk:paper-default@1.0.0",
max_gross_exposure=1.0,
max_single_asset_weight=0.6,
max_positions=10,
),
code_revision="c" * 40,
created_at=datetime(2026, 1, 9, 1, tzinfo=UTC),
top_k=2,
execution_price_field="open",
valuation_price_field="close",
execution_config=_execution_config(),
)