BuySideBenchAgentic portfolio management
Release 1.0
Published results

BuySideBench methodology

How BuySideBench evaluates portfolio decisions

Each task specifies an investment decision, supplies point-in-time evidence, and requires a concrete portfolio action. The system chooses its own research and analytical methods.

01

Anatomy of a task

The task specifies the investor's objective, permitted instruments, decision date, costs, and required output. It does not prescribe an estimator, optimizer, research sequence, or tool stack.

Briefing supplied to the system

  • Investor mandateThe decision requested, the objective, and the horizon, in the investor's units
  • Investor rulesHard limits and budgets; withheld until asked for in the interactive condition
  • Point-in-time evidenceData tables that end at the decision date; nothing from afterwards
  • Costs and action formatTrading costs and the exact schema the submitted action must follow

Evaluator data

  • Task modelThe latent structure and future innovations that generated the evidence
  • Private scenario setThe same plausible future outcomes are used to evaluate every system
  • Feasibility check and scorerHard-constraint check followed by the task's economic objective
  • Reference and ceilingThe sensible default maps to 0; the best attainable feasible action maps to 1

Investor-held facts

Position limits, sleeve ranges, budgets

No amount of analysis can recover them from market data. In the interactive condition they are missing from the briefing until the system asks.

Latent market structure

Which signals persist, which states are active

Must be inferred from the supplied evidence. Asking the investor cannot reveal it; the simulated investor does not know it either.

Evaluator-private futures

The private scenario set and task model

Never exposed to the system. Used after submission to evaluate the action across plausible futures.

Every task separates three kinds of information: facts only the investor holds, structure only the evidence can reveal, and futures only the evaluator ever sees.

Held constant

  • The decision the investor needs
  • The point-in-time evidence
  • The permitted action space
  • The economic scoring rule

Chosen by the system

  • Which evidence deserves weight
  • How uncertainty should be estimated
  • Which candidate actions should be compared
  • How the final decision should be implemented

02

Where the evidence comes from

Each task uses synthetic or semi-synthetic evidence, identified in the task record. In both cases the evaluator knows the data-generating process and can calculate the strongest attainable action.

Fully synthetic evidence

Fund a defensive overlay for a concentrated risky sleeve

One holding from the agent-facing return history

Every generating parameter is a documented design choice, so ground truth is exact and the private futures are drawn from the same coherent model that produced the history.

Semi-synthetic evidence

Replicate a target strategy return stream with liquid proxies

One liquid proxy from the agent-facing return history

Real component: Real Federal Reserve H.10 and H.15 daily series, October 2023 to March 2026, transformed into proxy returns.

Controlled component: A controlled synthetic target strategy overlaid on the real proxies, so the quantity to be replicated has known structure.

Both series come from the evidence supplied to the system. Because the evaluator knows the data-generating process, it can score without using future information and can calculate the strongest attainable action.

03

From historical evidence to future outcomes

A single market path is a noisy judge of a decision made under uncertainty. The historical evidence and the private evaluation scenarios come from the same task model, with different future innovations across scenarios.

History available to the systemPlausible future range, 10th to 90th percentileMedian future
Replicate a target strategy return stream with liquid proxies. The solid line is the target strategy, cumulative net return the system can inspect; it ends at the decision date. The fan is an illustrative panel of 240 scenarios drawn separately from the private scoring set. Units: cumulative net return, decimal.

04

Evaluating an action across scenarios

The scorer applies the submitted action to every scenario in the private evaluation set. The task specifies the relevant statistic: a mean, tail mean, worst case, or tracking error.

Mean value

expected value

Score reads +0.0038 , the mean outcome. One action: downside resilient, 240 illustrative scenarios.

Tail value

lower-tail mean (worst 5%)

Score reads -0.1543 , the mean of the shaded worst 5%. One action: diversified hedge, 240 illustrative scenarios.

Worst case

worst case

Score reads -0.2423 , the single worst outcome. One action: balanced collar, 64 illustrative scenarios.

Tracking error

tracking error (annualized RMSE)

Score reads +0.0706 , the annualized tracking RMSE. One action: rolling window ols projected, 240 illustrative scenarios.

Each panel holds one submitted action fixed and replays it across every scenario of an illustrative panel; the vertical line marks the single number the mandate reads from that distribution. Units differ because each mandate scores in its own economic units. Each outcome is that scenario's annualized RMS active return; the score aggregates them into one annualized tracking RMSE.

The scorer measures the economic objective defined for the task, such as expected return after costs, drawdown-aware utility, expected shortfall improvement, implementation shortfall, or tracking error. It calculates the expectation directly when a formula is available and otherwise averages across the evaluation scenarios.

05

Constraint checks and score normalization

The scorer checks hard constraints before measuring value. Feasible actions are compared with a sensible default and the strongest attainable action under the task model.

Constraint check

Every submitted action is checked against the investor's hard rules before any value is computed.

Feasible

The scorer measures how much of the available economic value the action captured.

Breaks a hard rule

The action is recorded as infeasible and reported as 0. The record distinguishes this from a feasible action that produces no improvement over the default.

The mapping is affine, so a submitted action sits at the same relative position on both axes. Scores below the default are negative and reported unclipped. The plotted dots are the two published results from the worked example on the overview page.

< 0

Worse than the task reference

0

Reference or infeasible action

0 to 1

Economic improvement over the reference

1

Strongest attainable action under the task model

The public comparison reports the median across tasks as its primary summary and retains task-level scores, feasibility results, submitted actions, and traces so the aggregate can be inspected.

06

The 14-task paired set

The table records the evidence source, capability tested, required action, objective, scoring method, and horizon for each task in the matched comparison.

AssignmentFamilyEvidenceMeasuresActionObjective aggregationScoringHorizon
Point-in-time corporate-event allocationSignal inferencesyntheticinference8 weightsexact expected valueclosed-form20 sessions
Cross-market information-propagation allocationSignal inferencesyntheticinference8 weightsexpected valuepanel, 19,584 scenarios20 sessions
Tactical allocation under uncertain market conditionsSignal inferencesyntheticinference9 weightsexpected valuepanel, 1,664 scenarios40 sessions
Market-neutral equity allocationPortfolio constructionsyntheticinference12 weightsexpected valuepanel, 16,384 scenarios40 sessions
Drawdown-aware strategic allocationPortfolio constructionsyntheticcompetent execution8 weightsexpected valuepanel, 16,384 scenarios126 sessions
Exclusion-aware regional completion portfolioPortfolio constructionsyntheticinference12 weightsexpected valuepanel, 16,384 scenarios63 sessions
Portfolio transition with competing research viewsPortfolio constructionsyntheticcompetent execution8 weightsexpected valuepanel, 16,384 scenarios60 sessions
Decide whether a market dislocation is investableRelative value & implementationsyntheticinference9 weightsexpected valuepanel, 16,384 scenarios63 sessions
Schedule an approved portfolio transition over five sessionsRelative value & implementationsyntheticcompetent execution30 weightsexpected value (minimized)panel, 265,728 scenarios5 sessions
Fund a defensive overlay for a concentrated risky sleeveRisk transfersyntheticcompetent execution9 weightslower-tail mean (worst 5%)panel, 65,536 scenarios21 sessions
Protect an equity sleeve with a bounded-cost option collarRisk transfersyntheticcompetent execution6 weightsworst casebounded support (exact)126 days
Build a liquid real-asset sleeve for observed persistent inflationRisk transfersyntheticinference10 weightsexpected valuepanel, 65,536 scenariosIndependent future state-active deployment sessions
Replicate a target strategy return stream with liquid proxiesReplication & data integritysemi-syntheticinference7 weightstracking error (annualized RMSE) (minimized)panel, 16,384 scenarios63 sessions
Set a reconciled allocation from raw point-in-time market extractsReplication & data integritysemi-syntheticinference8 weightsexpected valuepanel, 98,304 scenarios63 sessions
The fourteen tasks in the matched comparison, rendered from each task's executable packet and design record. Each row states what evidence the system receives, the action it must submit, and how that action is evaluated.

07

How task difficulty is established

For an inference task to be included, the evidence must support the required inference and a better estimate must lead to a better portfolio decision.

  1. 01

    Can the unknown be recovered?

    The evidence must be capable of revealing what the agent needs to learn.

  2. 02

    Is there enough signal?

    The observations must contain useful information, not just a theoretically relevant field.

  3. 03

    Would it change the portfolio?

    Learning the unknown must improve the best decision after costs and constraints.

  4. 04

    Does the advantage persist?

    The informed action must remain better across plausible futures, not just one convenient path.

A technical-looking task may still leave little room for research quality to matter. Constraints can narrow the optimum to one possible answer, market prices can dominate the research signal, or a simple heuristic can already be close to the strongest attainable action.

08

Two matched information conditions

The interactive benchmark and its all-information control test the same investment problem. Only the timing of investor-specific information changes.

All information supplied

Single-turn control

  • Investor rules are supplied in the initial briefing
  • The agent can begin research immediately
  • Measures decision quality with nothing withheld

Interactive

Multi-turn benchmark

  • Investor rules are absent from the initial briefing
  • The agent must identify which facts it needs and ask
  • The simulated investor answers only questions covered by its private document

All 14 tasks are evaluated in both conditions. The investment problem and scoring remain the same; only the timing of investor-specific information changes.