BuySideBenchAgentic portfolio management
Release 1.0
Published results

BuySideBench 1.0 · Itoflow Research

A benchmark for quantitative portfolio decisions

BuySideBench gives an AI system an investor request, point-in-time market data, trading costs, and portfolio constraints. The score is based on the submitted portfolio action.

What the benchmark tests

Research under uncertainty and investor constraints

The tasks require systems to interpret evidence, account for costs and constraints, and make a portfolio decision without knowing which market outcome will occur.

History available to the systemPlausible future range, 10th to 90th percentileMedian future

Itoflow researches investment questions and produces portfolio decisions subject to investor constraints. BuySideBench tests that work in two settings: complete initial instructions and assignments where the system must ask for missing investor constraints. Itoflow is evaluated on the same tasks and by the same scorer as the other systems shown below.

A task in practice

The system had enough evidence, but not enough investor information.

One assignment asks the system to improve an eight-equity portfolio after new earnings and guidance. It has the data needed to research the companies, but the brief does not reveal what portfolio the investor will accept. That distinction is the center of the benchmark.

What the system could research

Signals

Point-in-time corporate-event allocation

Improve an eight-equity sleeve after the latest earnings and guidance updates using only information available at the decision time.

Current weights, four years of issuer returns, timestamped event updates, market-factor returns, and a pre-outcome peer classification for 8 investable and 24 reference issuers.

5 data tables2,616 observationsNothing after the decision date

A preview of the supplied event data

4 of 520 rows · event_updates.csv

IssuerAvailable to systemEarnings surpriseGuidance change
EQ0131 Mar 2026-1.0+1.0
EQ0231 Mar 2026+1.0-1.0
EQ0331 Mar 2026-1.0+1.0
EQ0431 Mar 2026+1.0+1.0

The full packet also includes current portfolio weights, daily issuer returns, market-factor returns, and peer groups.

What only the investor could answer

None of those tables says how concentrated the investor is willing to be. Three rules needed to construct the portfolio are absent from the brief:

Maximum weight in any single asset: 30%Minimum number of held positions: fourMinimum funded weight: 8%

Better analysis cannot recover a client preference from market data. Before allocating, the system has to recognize the gap and decide whether to ask or assume.

The paths diverged here

Both systems completed the quantitative work. Only one paused to recover the investor's rules.

Asked, then acted

Itoflow · GPT-5.6-Sol high

0.550

Task score

Stopped before portfolio construction and asked one bundled question for the concentration limit, minimum holdings, and minimum funded weight.

Investor exchange

Asked: Should I use a conservative envelope based on the supplied portfolio, or will you provide the exact limits?

Client: Maximum weight per asset 30%; minimum holdings four; minimum funded weight 8%.

Received all three values, then submitted a feasible four-position allocation.

Submitted portfolio

EQ01
0%
EQ02
30%
EQ03
30%
EQ04
0%
EQ05
30%
EQ06
0%
EQ07
10%
EQ08
0%
Feasible action

Inferred, then acted

Codex · GPT-5.6-Sol high

0.187

Task score

Found the omitted limits but did not ask. It inferred a conservative range from the current portfolio and retained all eight positions.

Investor exchange

No question was sent to the client before submission.

Submitted a feasible but more conservative allocation without receiving the client limits.

Submitted portfolio

EQ01
9.375%
EQ02
15.625%
EQ03
15.625%
EQ04
15.625%
EQ05
15.625%
EQ06
9.375%
EQ07
9.375%
EQ08
9.375%
Feasible action

Both portfolios were feasible. Itoflow asked for the missing rules before allocating and scored 0.550. Codex inferred the rules and scored 0.187.

One example shows the mechanism. The full study asks whether the same advantage survives across different portfolio decisions and different kinds of missing investor information.

From one case to a controlled comparison

Same decision, two information conditions

The worked example captures one behavior. To isolate that behavior from general quantitative ability, every task is evaluated twice. The investment problem is unchanged; only the timing of investor-specific information differs.

Interactive

The multi-turn benchmark

The investor's rules are missing from the briefing. The system must recognize what it needs and ask before submitting an action.

All information supplied

The single-turn control

The same investor rules appear in the briefing from the start. The system can begin research immediately.

Same taskSame evidenceSame action spaceSame scorerSame possible futures
14
tasks in the benchmark
3
agent systems, 7 configurations
2
information conditions
196
completed evaluations

What the comparison revealed

Scores fell when systems had to recover missing information

With every investor constraint supplied in the initial briefing, the best Itoflow configuration reached a median of 0.778 and the best model-matched Codex configuration 0.774. All configurations scored lower when they had to identify and request the missing constraints.

Itoflow and Codex both run GPT-5.6-Sol, so that pair is the model-matched comparison; Claude Code runs Opus 5. Each system uses its own tools, code, and research process. The scorer judges only the submitted action.

Results as of 1 August 2026

Median score by information condition

All information suppliedInteractive

Itoflow

GPT-5.6-Luna · XHigh reasoning

0.7380.689change -0.048

Itoflow

GPT-5.6-Sol · Low reasoning

0.6850.663change -0.023

Itoflow

GPT-5.6-Sol · High reasoning

0.7780.489change -0.290

Codex

GPT-5.6-Sol · High reasoning

0.7510.349change -0.402

Claude Code

Opus 5 · High reasoning

0.6040.306change -0.297

Codex

GPT-5.6-Sol · Low reasoning

0.7740.228change -0.546

Claude Code

Opus 5 · Low reasoning

0.4920.107change -0.385

Feasible actions in the interactive condition

When a system assumes missing investor constraints, it may submit a portfolio that breaks a hard limit. Those actions remain in the results and receive a score of 0. With all information supplied, 96 of 98 submitted actions were feasible.

Itoflow

GPT-5.6-Luna · XHigh reasoning

14/14 feasible

Itoflow

GPT-5.6-Sol · Low reasoning

13/14 feasible

Itoflow

GPT-5.6-Sol · High reasoning

12/14 feasible

Claude Code

Opus 5 · High reasoning

10/14 feasible

Claude Code

Opus 5 · Low reasoning

10/14 feasible

Codex

GPT-5.6-Sol · High reasoning

8/14 feasible

Codex

GPT-5.6-Sol · Low reasoning

7/14 feasible

Interactive

The initial briefing omits investor constraints that cannot be inferred from market data. The system must ask for them before submitting an action.

Median is the primary summary because a few extreme task scores can move the mean. Feasibility is reported because a numerically attractive action does not count if it breaks the investor's rules.

SystemConfigurationMedianMeanTasks completedFeasible actions
ItoflowGuardian: GPT-5.6-Luna xhighGPT-5.6-LunaXHigh0.6890.56214/1414/14
ItoflowGuardian: GPT-5.6-Sol highGPT-5.6-SolLow0.6630.57514/1413/14
ItoflowGuardian: GPT-5.6-Sol highGPT-5.6-SolHigh0.4890.43114/1412/14
CodexGPT-5.6-SolHigh0.3490.35514/148/14
Claude CodeOpus 5High0.3060.33114/1410/14
CodexGPT-5.6-SolLow0.2280.34414/147/14
Claude CodeOpus 5Low0.1070.23814/1410/14

A score of 0 matches the task's sensible default, 1 is the strongest attainable action under the task model, and negative values fall below the default. Actions that break a hard investor rule also receive 0.

Itoflow rows include its Guardian, a built-in reviewer that checks the research and the final action before submission; the benchmark awards it no points. In the research paper, the interactive condition is the multi-turn benchmark and the all-information condition is the single-turn control.

Release 1.0 reports one completed evaluation per task and configuration. Multi-run confidence intervals are in progress and will be added to these charts.

Behind the medians

All 196 scores

The aggregate result is not uniform. This matrix shows where each system gained from complete information, where it remained robust without it, and where a submitted action broke an investor rule.

How to read a score
Negative is worse than the task's sensible default, such as leaving the existing portfolio unchanged. 0 is no economic improvement over that default. Values toward 1 approach the strongest attainable decision under the task's own model.
What × means
The system completed the assignment, but its submitted action broke a hard investor constraint, such as a position limit. The action is preserved in the record and scored 0.
Below the defaultMatches the defaultToward the best attainableBroke a hard constraint; scored 0

Swipe horizontally to compare configurations.

AssignmentInteractiveAll information supplied
ItoflowHighItoflowLowItoflowXHighCodexHighCodexLowClaudeHighClaudeLowItoflowHighItoflowLowItoflowXHighCodexHighCodexLowClaudeHighClaudeLow
Point-in-time corporate-event allocationSignals0.5500.7010.7010.1870.7010.2650.7010.7010.8940.3760.454
Cross-market information-propagation allocationSignals0.0250.000-0.3660.538-0.0100.082-0.582-0.582-0.5820.2880.2880.1960.000
Tactical allocation under uncertain market conditionsSignals0.0000.1920.1320.3370.2120.195-0.1050.2120.1160.038
Market-neutral equity allocationPortfolio construction0.9150.8990.9070.8520.8620.8300.7480.856-0.0510.9070.8340.9070.8650.859
Drawdown-aware strategic allocationPortfolio construction0.7830.8560.9540.7310.457-0.0090.9530.9120.9530.9530.9730.666-0.202
Exclusion-aware regional completion portfolioPortfolio construction0.6780.1760.6780.6550.4560.3180.3950.1760.6780.6780.6780.6780.1960.267
Portfolio transition with competing research viewsPortfolio construction0.9350.8930.9350.9280.9280.7790.9280.9230.9350.9350.9350.9350.9850.923
Decide whether a market dislocation is investableRelative value0.0000.2570.3200.5360.6450.3370.3430.5450.6640.4020.6490.3470.5410.529
Schedule an approved transition over five sessionsExecution0.9940.4720.4750.5110.8160.7060.6750.9920.9270.7750.4600.4910.9420.890
Fund a defensive overlay for a concentrated sleeveRisk transfer0.0001.0001.0001.0001.0001.0001.0001.0001.0000.973
Protect an equity sleeve with a bounded-cost collarRisk transfer0.0001.0001.0001.0001.0001.0001.0001.0001.000
Build a liquid real-asset sleeve for persistent inflationRisk transfer0.7330.8701.0000.7330.7330.8700.8700.8700.8700.8700.9140.965
Replicate a target strategy with liquid proxiesInstitutional0.6250.5740.5730.5730.2950.6930.6930.5740.6070.5740.390
Set an allocation from raw point-in-time extractsInstitutional0.4270.301-0.306-0.6890.453-0.6140.9080.800-0.944-0.270-0.269

Cell color shows score direction and magnitude; the numbers are the exact unclipped scores. A few hedging assignments sit near 1.000 for most configurations in the all-information condition. They chiefly test correct execution, and the methodology covers how much room each task leaves beyond it.

Inspect the underlying work

Open any assignment from request to score

The matrix is the summary. The task explorer below exposes what each system received, which investor rules changed between conditions, what action it had to submit, and how that action was judged.

Signals

Point-in-time corporate-event allocation

signal.event_response

Interactive view

Investor request

Improve an eight-equity sleeve after the latest earnings and guidance updates using only information available at the decision time.

Evidence shown

Current weights, four years of issuer returns, timestamped event updates, market-factor returns, and a pre-outcome peer classification for 8 investable and 24 reference issuers.

5 public tables2,616 observationsPoint-in-time data

Required action and economic objective

Eight long-only opening weights.

Expected incremental terminal return versus the supplied portfolio, net of opening turnover cost.

Observed approaches on this task

Itoflow · GPT-5.6-Sol high
0.550

Stopped before portfolio construction and asked one bundled question for the concentration limit, minimum holdings, and minimum funded weight.

Investor exchange

Asked: Should I use a conservative envelope based on the supplied portfolio, or will you provide the exact limits?

Client: Maximum weight per asset 30%; minimum holdings four; minimum funded weight 8%.

Received all three values, then submitted a feasible four-position allocation.

Feasible action
Codex · GPT-5.6-Sol high
0.187

Found the omitted limits but did not ask. It inferred a conservative range from the current portfolio and retained all eight positions.

Investor exchange

No question was sent to the client before submission.

Submitted a feasible but more conservative allocation without receiving the client limits.

Feasible action

Application to Itoflow

Why missing investor constraints matter

Investors do not always provide every constraint upfront. Itoflow is built to ask for decision-relevant values before proposing a portfolio. BuySideBench tests that behavior directly.

Each result scores the submitted action, not Itoflow's internal workflow, evidence system, or review machinery.

Run your system against BuySideBench

The next result on this page should not have to come from us. Model developers, researchers, and investment firms can request access while the public release is prepared.

The public repository will include the assignments, harness and scorer interfaces, and the records behind the results on this page. The research paper will document task construction and evaluation in full. Until those materials are published, we are onboarding participants directly.

Public GitHub
Coming soon
Research paper
Coming soon
Early access
hello@itoflow.ai
    BuySideBench | Itoflow