146 lines
9.5 KiB
Markdown
146 lines
9.5 KiB
Markdown
# Retrospective computation contracts v2 (unreleased)
|
|
|
|
This pure, storage-neutral compatibility path consumes the separate data-contract
|
|
major 2.0.0. It does not migrate, reinterpret or relax the accepted v1 contracts.
|
|
No financial formula, execution simulation, dependency lock, production database,
|
|
publisher or live/paper-order interface changes here. Package version is unchanged;
|
|
the new contract major is not a package release or deployment.
|
|
|
|
## Explicit public boundaries
|
|
|
|
| Module | Public types/builders | Changed wire identity |
|
|
| --- | --- | --- |
|
|
| `retrospective_data_contracts` | `RetrospectiveSnapshotEnvelope`, `RetrospectiveFoundationEnvelope` | `rhdsv2`, `rhdfv2`; consume RP-owned 2.0.0 data semantics |
|
|
| `retrospective_factor_contracts` | `RetrospectiveFactorSetRef`, typed input/view/causation bindings, `ResolvedRetrospectiveView` | `rhfactorsetv2` |
|
|
| `retrospective_backtest_contracts` | `RetrospectiveBacktestRunRef` | `rhbacktestrunv2` |
|
|
| `retrospective_artifact_contracts` | `RetrospectiveBacktestEvidenceManifest`, `RetrospectivePerformanceEvidence` and their builders | `rhbacktestevidencev2`, `rhperformancev2` |
|
|
| `retrospective_portfolio_risk_contracts` | `RetrospectivePortfolioTarget`, `RetrospectivePortfolioDecision`, `RetrospectiveRiskAssessment`; receipt-digest, decision and assessment builders | `rhportfoliotargetv2`, `rhportfoliodecisionv2`, `rhriskassessmentv2` |
|
|
|
|
These are separate types and domain-separated content identities. There is no
|
|
automatic v1-to-v2 cast. Unknown schema versions and fields are rejected. The
|
|
performance wire keeps its named schema `researchhub.performance-evidence.v2`;
|
|
the other new computation contracts use `schema_version: 2.0.0`.
|
|
|
|
FactorDefinition, factor-output byte references, output quality/coverage,
|
|
ConstraintSetV1, FreshnessPolicy, ComputationReceipt, CovarianceSnapshot, financial
|
|
algorithms, performance metric/methodology IDs and the nine ResearchRunArtifact
|
|
tables keep their existing semantics. The table schema remains **1.1.0**. Reusing
|
|
these neutral primitives does not make a new-major upstream reference v1-compatible.
|
|
|
|
## Two clocks, not backdated evidence
|
|
|
|
Every new result fixes `usage=retrospective_research` and
|
|
`historical_availability=not_established`. A business date describes the historical
|
|
period being researched. Observation, publication, evaluation, artifact availability,
|
|
target creation and computation describe actual events, and must not be backdated.
|
|
Public v2 instants require UTC `Z` with at most six fractional digits.
|
|
|
|
`observation_cutoff` and chunk `observed_by` are upper-bound observations. They are
|
|
not the earliest public knowledge time or a PIT cutoff. Unknown earliest knowledge
|
|
stays unknown; a supplied knowledge-evidence digest is not authenticated by parsing.
|
|
Foundation observation sequences describe retained revisions, not complete original
|
|
history. Selected view routes, calendars, corporate-action coverage and lineage
|
|
must close exactly within the supplied Foundation.
|
|
|
|
Required actual order for factor/backtest evidence is:
|
|
|
|
1. Foundation publication <= factor evaluation <= factor computation <= factor availability.
|
|
2. Factor availability <= backtest evaluation <= artifact start <= artifact finish
|
|
<= backtest computation <= artifact availability.
|
|
3. Artifact availability <= target creation <= portfolio computation <= risk computation.
|
|
|
|
RetrospectivePortfolioTarget has a historical `effective_at` and a distinct actual
|
|
`created_at`. PortfolioDecision carries both plus actual `computed_at`. Covariance
|
|
window end <= covariance as-of date <= the historical effective date; covariance
|
|
maximum age is measured against that historical date. Manifest maximum age is
|
|
measured against **actual** portfolio and risk computation separately. Passing one
|
|
age check cannot substitute for the other. Generic v1 receipt timestamps retain
|
|
their original normalization; binding compares parsed actual instants.
|
|
|
|
## Materialized bytes and reference-only reads
|
|
|
|
Snapshot decoding checks structure, all six blocking-quality declarations,
|
|
qualification/time ordering, observation receipts and identities.
|
|
`verify_materialized_records` additionally checks supplied chunks, per-chunk and
|
|
aggregate content, counts, dimensions, effective ranges and macro effective instants.
|
|
Provider/physical paths are forbidden in public metadata and materialized records.
|
|
|
|
Factor creation requires actual snapshot chunks, selected view schema/content bytes,
|
|
and factor-output schema/content bytes. Definition inputs, view availability,
|
|
Foundation ancestry and computed digests must close. Reference-only deserialization
|
|
is allowed for display/inspection, but input/output validation flags are derived from
|
|
supplied bytes, are not serialized claims, and must be re-established for new
|
|
computation. Backtest creation requires a factor whose payloads were revalidated.
|
|
Reference decoding cannot turn an unverified factor into an admitted compute input.
|
|
|
|
Backtest manifest decoding rebuilds evidence from the supplied typed run and all
|
|
nine actual artifact tables. It checks table/run/config/strategy bindings and time
|
|
ordering. Portfolio composition revalidates those retained tables again, rather
|
|
than trusting a serialized manifest or mutable Python context. A table digest proves
|
|
content binding, not that those tables were produced by the claimed computation.
|
|
|
|
All content-addressed IDs exclude their own ID field and bind the remainder of the
|
|
closed payload. Data/factor/backtest/manifest JSON retains the strict data profile
|
|
(no JSON floating-point numbers; financial record decimals are strings). Performance
|
|
and S4 preserve the existing finite numeric JSON profile: finite floats, safe ints,
|
|
exact booleans, sorted keys, compact separators, UTF-8. Duplicate keys, NaN,
|
|
Infinity, noncanonical JSON and extra fields are rejected. Wire revalidation uses
|
|
type-sensitive comparisons, including `true` versus `1`. Serializers return
|
|
detached copies; internal public maps are immutable.
|
|
|
|
## Replay, receipts and risk
|
|
|
|
Backtest v2 replay specification binds immutable input identities, selected calendar
|
|
and actions, strategy/execution/cost versions and digests, configuration, code,
|
|
environment lock and random seed. It excludes **both actual evaluation and actual
|
|
computation time**. These actual times remain in each run's identity. A replay must
|
|
retain the same replay specification, append its unique full ancestry, increment
|
|
attempt by one, and have parent computation < new actual evaluation <= computation.
|
|
This explicit new-major rule allows a later genuine replay without pretending its
|
|
evaluation happened at the parent's clock time.
|
|
|
|
Portfolio computation-input v2 binds the full run and manifest document digests,
|
|
new target (including both clocks), objective/model versions and digests, declared
|
|
expected returns/covariance/scenario inputs, freshness policy and prior weights.
|
|
The receipt separately binds that input, constraints and recomputed outputs/residuals.
|
|
Targets and prior holdings must use selected logical instrument IDs, not ad-hoc
|
|
symbol matches. Failed/fallback receipts and any actual constraint residual are
|
|
rejected, even if a solver declares convergence within a permissive tolerance.
|
|
|
|
Risk uses the existing labelled Euler decomposition exactly once. Its result binds
|
|
the supplied matrix content plus covariance method, bounded estimation window,
|
|
observation count, lookback, missing policy, annualization, source dataset/input,
|
|
model/budgets/groups and actual computation time. It checks exact labels, finite
|
|
symmetry, covariance-source binding and both freshness clocks. Non-PSD,
|
|
non-positive portfolio variance or non-closed contributions produce an unavailable,
|
|
unqualified result. A budget breach is a ready but unqualified calculation result.
|
|
`qualified=true` means only that these calculation checks passed. Every result
|
|
remains `decision_eligible=false`, `execution_validation=not_validated`; no portfolio
|
|
approval, maker-checker, publication, paper or live permission is granted here.
|
|
|
|
## Trust, ownership and test evidence
|
|
|
|
Pure builders accept declarations. Hashes, typed objects, model names, successful
|
|
constraint checks and synthetic fixtures do **not** authenticate data or compute
|
|
producers. Trusted owner-version bindings and receipt/qualification/view/clock
|
|
admission ports remain mandatory. RP owns governance and presentation; Research
|
|
Results owns publication. QE supplies validated calculation facts only.
|
|
|
|
The two data fixtures are public RP candidate vectors from
|
|
`7da27e5bd33dc6d06f2c7c60f47029111156293a` (PR #100). EDB producer candidate
|
|
`88433dfc9d865ef782465498cdf9454c73920abd` (PR #13) is not a runtime dependency or
|
|
accepted owner binding. Acceptance/review gates remain separate from local tests.
|
|
|
|
`tests/fixtures/retrospective-computation-v2.golden.json` freezes newly constructed
|
|
synthetic factor/run/manifest/performance/portfolio/risk payloads and their inputs.
|
|
Its artifact matrices are separate synthetic envelope-test inputs: the one-day
|
|
public data fixture is **not** claimed to have produced the four-day artifact.
|
|
The vector is not an end-to-end data/computation provenance proof or real-data run.
|
|
Its deterministic IDs are contract-regression evidence, not admitted source facts.
|
|
|
|
Focused tests cover v2 goldens, mutation and strict JSON, bytes versus references,
|
|
two-clock freshness, replay ancestry, exact table bindings, constraints, receipts,
|
|
covariance provenance and numerical findings. Existing v1 tests must also pass.
|
|
Rollback is disabling the explicit v2 entry path while retaining v1 and original
|
|
immutable results; never retag old results or silently downgrade failed v2 admission.
|