Investor-held facts
Position limits, sleeve ranges, budgets
No amount of analysis can recover them from market data. In the interactive condition they are missing from the briefing until the system asks.
BuySideBench methodology
Each task specifies an investment decision, supplies point-in-time evidence, and requires a concrete portfolio action. The system chooses its own research and analytical methods.
01
The task specifies the investor's objective, permitted instruments, decision date, costs, and required output. It does not prescribe an estimator, optimizer, research sequence, or tool stack.
Briefing supplied to the system
Evaluator data
Position limits, sleeve ranges, budgets
No amount of analysis can recover them from market data. In the interactive condition they are missing from the briefing until the system asks.
Which signals persist, which states are active
Must be inferred from the supplied evidence. Asking the investor cannot reveal it; the simulated investor does not know it either.
The private scenario set and task model
Never exposed to the system. Used after submission to evaluate the action across plausible futures.
02
Each task uses synthetic or semi-synthetic evidence, identified in the task record. In both cases the evaluator knows the data-generating process and can calculate the strongest attainable action.
Fully synthetic evidence
One holding from the agent-facing return history
Every generating parameter is a documented design choice, so ground truth is exact and the private futures are drawn from the same coherent model that produced the history.
Semi-synthetic evidence
One liquid proxy from the agent-facing return history
Real component: Real Federal Reserve H.10 and H.15 daily series, October 2023 to March 2026, transformed into proxy returns.
Controlled component: A controlled synthetic target strategy overlaid on the real proxies, so the quantity to be replicated has known structure.
03
A single market path is a noisy judge of a decision made under uncertainty. The historical evidence and the private evaluation scenarios come from the same task model, with different future innovations across scenarios.
04
The scorer applies the submitted action to every scenario in the private evaluation set. The task specifies the relevant statistic: a mean, tail mean, worst case, or tracking error.
expected value
Score reads +0.0038 , the mean outcome. One action: downside resilient, 240 illustrative scenarios.
lower-tail mean (worst 5%)
Score reads -0.1543 , the mean of the shaded worst 5%. One action: diversified hedge, 240 illustrative scenarios.
worst case
Score reads -0.2423 , the single worst outcome. One action: balanced collar, 64 illustrative scenarios.
tracking error (annualized RMSE)
Score reads +0.0706 , the annualized tracking RMSE. One action: rolling window ols projected, 240 illustrative scenarios.
The scorer measures the economic objective defined for the task, such as expected return after costs, drawdown-aware utility, expected shortfall improvement, implementation shortfall, or tracking error. It calculates the expectation directly when a formula is available and otherwise averages across the evaluation scenarios.
05
The scorer checks hard constraints before measuring value. Feasible actions are compared with a sensible default and the strongest attainable action under the task model.
Constraint check
Every submitted action is checked against the investor's hard rules before any value is computed.
Feasible
The scorer measures how much of the available economic value the action captured.
Breaks a hard rule
The action is recorded as infeasible and reported as 0. The record distinguishes this from a feasible action that produces no improvement over the default.
Sensible default
The task's economic starting point
Normalized score: 0
Submitted action
The value captured by the system
Measured in the task's own units
Best feasible action
The strongest attainable decision
Normalized score: 1
score = (submitted value - default value) / (best feasible value - default value)
< 0
Worse than the task reference
0
Reference or infeasible action
0 to 1
Economic improvement over the reference
1
Strongest attainable action under the task model
The public comparison reports the median across tasks as its primary summary and retains task-level scores, feasibility results, submitted actions, and traces so the aggregate can be inspected.
06
The table records the evidence source, capability tested, required action, objective, scoring method, and horizon for each task in the matched comparison.
| Assignment | Family | Evidence | Measures | Action | Objective aggregation | Scoring | Horizon |
|---|---|---|---|---|---|---|---|
| Point-in-time corporate-event allocation | Signal inference | synthetic | inference | 8 weights | exact expected value | closed-form | 20 sessions |
| Cross-market information-propagation allocation | Signal inference | synthetic | inference | 8 weights | expected value | panel, 19,584 scenarios | 20 sessions |
| Tactical allocation under uncertain market conditions | Signal inference | synthetic | inference | 9 weights | expected value | panel, 1,664 scenarios | 40 sessions |
| Market-neutral equity allocation | Portfolio construction | synthetic | inference | 12 weights | expected value | panel, 16,384 scenarios | 40 sessions |
| Drawdown-aware strategic allocation | Portfolio construction | synthetic | competent execution | 8 weights | expected value | panel, 16,384 scenarios | 126 sessions |
| Exclusion-aware regional completion portfolio | Portfolio construction | synthetic | inference | 12 weights | expected value | panel, 16,384 scenarios | 63 sessions |
| Portfolio transition with competing research views | Portfolio construction | synthetic | competent execution | 8 weights | expected value | panel, 16,384 scenarios | 60 sessions |
| Decide whether a market dislocation is investable | Relative value & implementation | synthetic | inference | 9 weights | expected value | panel, 16,384 scenarios | 63 sessions |
| Schedule an approved portfolio transition over five sessions | Relative value & implementation | synthetic | competent execution | 30 weights | expected value (minimized) | panel, 265,728 scenarios | 5 sessions |
| Fund a defensive overlay for a concentrated risky sleeve | Risk transfer | synthetic | competent execution | 9 weights | lower-tail mean (worst 5%) | panel, 65,536 scenarios | 21 sessions |
| Protect an equity sleeve with a bounded-cost option collar | Risk transfer | synthetic | competent execution | 6 weights | worst case | bounded support (exact) | 126 days |
| Build a liquid real-asset sleeve for observed persistent inflation | Risk transfer | synthetic | inference | 10 weights | expected value | panel, 65,536 scenarios | Independent future state-active deployment sessions |
| Replicate a target strategy return stream with liquid proxies | Replication & data integrity | semi-synthetic | inference | 7 weights | tracking error (annualized RMSE) (minimized) | panel, 16,384 scenarios | 63 sessions |
| Set a reconciled allocation from raw point-in-time market extracts | Replication & data integrity | semi-synthetic | inference | 8 weights | expected value | panel, 98,304 scenarios | 63 sessions |
07
For an inference task to be included, the evidence must support the required inference and a better estimate must lead to a better portfolio decision.
The evidence must be capable of revealing what the agent needs to learn.
The observations must contain useful information, not just a theoretically relevant field.
Learning the unknown must improve the best decision after costs and constraints.
The informed action must remain better across plausible futures, not just one convenient path.
A technical-looking task may still leave little room for research quality to matter. Constraints can narrow the optimum to one possible answer, market prices can dominate the research signal, or a simple heuristic can already be close to the strongest attainable action.
08
The interactive benchmark and its all-information control test the same investment problem. Only the timing of investor-specific information changes.
Single-turn control
Multi-turn benchmark
All 14 tasks are evaluated in both conditions. The investment problem and scoring remain the same; only the timing of investor-specific information changes.