tradefloor

io.github.simoncoombesv0.9.0更新于 Oct 3, 2026

Simulated stock markets that react to your orders: run, fork and score trading agents.

已验证STDIO仅桌面AI & MLFinanceData & Analytics

概览

AI 生成的概览

运行确定性的模拟股票市场,让助手测试、分叉并评分交易策略与智能体。

功能
tradefloor 是一个带限价订单簿的市场模拟器:给定随机种子和一组公司,它会推进价格、成交和每日经济,订单与订单簿深度撮合,因此交易会推动价格。通过 MCP,它把策略、股票池和情景作为数据暴露出来,助手可以跨多个种子评估或排名交易智能体、分叉一个运行中的市场并在某一分支中改变一个条件,并读取价格变动背后的因子级解释。结果自带其适用范围的说明。
适用场景
当你需要对交易行为做反事实实验时使用:换一种策略、加息或流动性危机下会发生什么,以及两个分支在哪里分道扬镳。它适合策略比较、智能体评分和研究流程,而不是实盘交易或真实市场数据。
运行要求
以 Python 包形式通过 stdio 在本地运行(tradefloor,需安装 mcp 附加组件,用 uvx 或 tradefloor-mcp 命令启动)。需要 CPython 3.11+,Linux、macOS 和 Windows 均有 wheel。未声明需要 API 密钥、账号或网络访问。仅限桌面端。
安装前请注意
它是模拟器而非预测:情景中的价格变动幅度只有真实规模的四分之一到一半,现实性结论也很有限(一年是经过认证的期限)。同一进程内的智能体可能泄漏状态或读取种子作弊,因此未受信任的智能体应放在独立进程中运行。没有佣金、借券费和止损单;除非设置 cash_interest,闲置现金没有收益。1.0 之前 API 可能变化。

安装

在 SourceWeft 中

  1. 打开 控制台中的 tradefloor,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

tradefloor

[determinism] [PyPI] [PyPI Downloads] [crates.io] [license: MIT OR Apache-2.0] [python 3.11+]

tradefloor is a market simulator you can run a strategy against. It has a Rust core and a Python API.

Give it a seed and a list of companies. It runs a market forward: prices, a limit order book, fills, and an economy that moves each day. Your orders match against the book's depth, so your trades move the price.

Real market data can't tell you what would have happened if you had traded differently, or what caused a move. tradefloor can, because it computed every price. You can fork a running market, change one thing in one branch (a rate rise, a liquidity crisis, a different agent), and measure where the two branches came apart. engine.truth() splits each move in the gap between a price and the model's fair value into eleven factors, and engine.explain(ticker, day) breaks down the move in the traded price, two records no historical dataset carries.

Documentation is at https://docs.tradefloor.dev.

Install

pip install tradefloor

There are wheels for Linux, macOS and Windows on CPython 3.11+, and no dependencies. The same engine is a Rust crate (cargo add tradefloor). Optional extras add the MCP server (tradefloor[mcp]), Arrow output (tradefloor[arrow]), the Gymnasium environment (tradefloor[rl]) and one extra per agent framework.

The API may change before 1.0. Model changes ship as new presets, so a market with no agent orders in it replays exactly on its named preset in later releases. tradefloor was called pretium until 0.5.0. Versions up to 0.4.3 still install under that name, and results recorded with them still replay.

A first run

python
import tradefloor as tf
universe = tf.Universe.random(40, seed=111)
spec = tf.StrategySpec.momentum(lookback_days=1.0, top_k=5)scores = tf.evaluate({"mine": spec}, seed=7, universe=universe, days=10)
scores["mine"].return_pct            # what it madescores["mine"].impact_bps            # what its own footprint costscores["mine"].strategy_fingerprint  # sha256, cite thisscores["mine"].errors                # each step that raised or was refusedscores["mine"].sharpe                # annualised, from the daily closesscores["mine"].time_in_market        # share of steps holding a position

That result comes from one random market, so it says as much about the seed as about the strategy. tf.rank runs many seeds and compares strategies with a paired sign test. Add tf.baselines.reference_agents() to the entrants and tf.versus_buy_and_hold(scores) reads each score against buy-and-hold on the same market.

A Python agent is any object with act(obs) that returns orders: a number of shares for a market order, tf.Limit(quantity, price) or tf.Cancel(). It sees a read-only view of the market and its own portfolio, and obs.history holds a daily bar per name. A bar's close is the day's last print. On pt-v20 the close then re-marks every name, so the next day starts 15 bp away at the median on a 20-name roster. docs/AGENTS.md covers what the view holds, how trades are charged, the framework adapters (OpenAI Agents SDK, PydanticAI, LangGraph, FinRobot) and how scoring works.

The demo

The examples are in this repository and not in the package, so clone it first:

git clone https://github.com/simoncoombes/tradefloorcd tradefloorpython examples/rate-shock/counterfactual.py

It runs an agent in a controlled market, checkpoints the world and forks it, raises rates by 200 bp in one branch, and compares what the same agent does next. It prints nine checks that the two branches started identical, the step at which the agent's behavior changed, and the two branches side by side. The run takes under five seconds of CPU and needs no keys and no network. The walkthrough is Your first counterfactual experiment.

Contents

engine.truth()why each price moved: eleven factors that sum to the mispricing's move, to 1e-16
engine.prints()how each trade price came about: the shock, and the order book depth that absorbed it
counterfactual TCAyour trading cost, from the same seed run with your orders and without them
tf.rankmany seeds, paired sign tests
RunManifestwhat a reader needs to replay a run, checked by reproduce()
World / comparefork a running experiment, change one variable, and measure where the two came apart
scenariosseven packaged shocks, and a file format for your own
MCP serverthirteen read-only tools for a coding agent, scenarios included
morea Gymnasium environment, Arrow output, checkpoints, SEC EDGAR data, simulated rate indices, a browser build

Drive it from an agent

pip install "tradefloor[mcp]"claude mcp add tradefloor -- tradefloor-mcp

tradefloor-mcp speaks MCP over stdio, and tradefloor mcp starts the same server. Strategies, universes and scenarios are data, so a tool argument cannot reach code. Each result carries its own caveats. See the MCP page.

Scenarios

python
engine = tf.Engine(seed=42, universe=universe)engine.run_days(20)                                # a shared history firstscenario = tf.Scenario.load("liquidity_crisis")    # ships with the package
control, stress = tf.branch(engine, 2)for day in range(80):    scenario.apply(stress, day)    ...                                  # run both branches

A scenario is a file of changes to the market and the assumptions behind them. Each change targets a field the engine reads, and the file keeps the shock apart from the knock-on effects you assume follow it:

tradefloor scenario show oil_price_spike
Exogenous shocks----------------------------------------------------------  day 50+            commodity.oil            x1.4
Assumed transmission----------------------------------------------------------  day 55..74 ramp    macro.inflation          +1.50pp  day 55+            macro.corporate_yield    +0.50pp

at in a scenario counts days from the first day it is applied, so on a branch it counts from the branch. Most packaged files first fire on day 50. To fire one on the first day after a fork, use scenario.starting_at(0), or world.apply(scenario, at=0) on a World. The gaps between its events stay the same, and the run's record keeps the packaged file's fingerprint and the days each event fired.

tradefloor does not predict what a war, an election, an oil shock or a recession will do to markets. You state the assumptions and it measures how an agent behaves under them. tradefloor scenario list names the seven packaged scenarios, and tradefloor scenario targets lists every field a scenario can change.

Reproducibility

The same seed gives the same market on every platform. tradefloor ships its own exp, log, pow, sin and cos, so the system's math library cannot change a result, and each release runs a fixed simulation on five platforms and stops if any result differs.

A shipped preset never changes, so a market with no agent orders in it replays exactly on its named preset in every later release. Each release checks that with a digest per preset. A run with agent orders in it replays exactly on the same release. Across releases the promise is narrower. 0.8.5 changed how an agent's fills reach the market, on every preset, so a traded run recorded before 0.8.5 matches up to its first trade and differs after it. The default preset is pt-v20, and any earlier one can be named:

python
eng = tf.Engine(seed=42, universe=u, model="pt-v10")

To let a reader rerun a result, publish its RunManifest. It records the version, preset, seed, universe, macro state and scenario, and reproduce() stops on a mismatch. A manifest checks the market and carries no score. Its result block holds the market's digest, the number of days and draws_consumed. tf.evaluate and tf.rank write no manifest, so a published score has to be rerun to be checked. docs/REPRODUCIBILITY.md has the full contract, including what a saved engine state promises when it is restored, and docs/SUPPORT.md says which release to pin for a long study.

Realism

tradefloor checks its market against real ones with three named sets of statistics, listed in docs/STATISTICS.md. On the default preset, pt-v20, all 19 statistics of the one-year table (volatility, fat tails, how much stocks move together, how far the VIX jumps after a fall) are inside the range real markets show over a year. All 14 graded statistics of the two-year panel are inside their two-year ranges. The long-run criteria are 40 rows over 21 years for pt-v20, covering crash depth, how long fear lasts, bear markets per decade, the 2008 and 2020 replays, the rate indices and the cost of size in the book. pt-v20 meets all 40.

Read those claims narrowly:

  • The 19 of 19 is a verdict on figures pooled over 30 seeds. One seed's year often misses some of its 14 shape statistics. On seeds 101 to 116, all 14 were in range on 5 of the 16, and one seed had 8 of 14. If you run one market per condition, read tf.envelope.intervals() for each statistic's spread across seeds.
  • A shape statistic's range is the median of 35 real one-year windows plus or minus 2.1 trimmed standard deviations, so passing one is weak evidence. Volatility clustering is one case. abs_return_acf1 reads 0.028, below every real 2015 to 2025 window (the lowest is 0.039), and it passes because its range reaches lower than those windows do.
  • The one-year table helped choose most of pt-v20's coefficients, so the held-out checks are the fresh seeds and the fresh set of companies the panel is repeated on.
  • One year is the certified horizon. Two years is graded on the two-year panel, and longer runs only by the long-run criteria. Every run on a roster opens at nearly the same VIX (17.66 on the certified roster), so the one-year figures describe years that start calm.
  • A driven scenario moves prices at a quarter to a half of the real size, in the right direction. Use a scenario to detect a response, and do not read its size as a forecast.
  • Volatility memory is weaker than real at every lag, about a quarter of real at lag 1. Nothing below the 65-minute step is calibrated.
  • An order sliced over a day costs far less than published studies find: 0.04 of a daily standard deviation for 10% of a day's volume in 36 slices, against 0.15 to 0.3. A schedule optimiser will overstate the value of trading slowly.
  • Your fills pay for the book depth they take, but that temporary impact barely reaches the printed prices. The lasting part is linear and fades, and no other trader adapts to you, so no liquidity spiral or predatory trading can arise.

tf.envelope.check(horizon_days=...) refuses a question that falls outside a measured limit. docs/REALISM.md has every number behind these claims and the full table of limits.

Before you publish a result

  • An agent scored on naming the factor behind each day's move gets an explanation_accuracy. On pt-v20 a constant answer scores 0.95 to 1.0, so quote explanation_edge, the accuracy minus that baseline, and never the accuracy alone.
  • Agents in one tf.evaluate or tf.rank call run one after another in one Python process, on one seed. An earlier agent can leave the price path in a class variable for a later one. An agent written to cheat can read the seed from the harness's frames through sys._getframe and run a copy of the market ahead. Nothing flags either. The read-only market view guards only against accidents, so run each agent you did not write in its own process, through the MCP server.
  • Every evaluate and rank run starts at day 0, so a rule that needs 20 days of prices sits out the first 20 while buy-and-hold is invested. Pass history_days=20 to run the market 20 days first with nobody trading.
  • In a World with several agents, orders placed at the same step execute in label order, alphabetical, for the whole run. Rotate the labels across runs when you compare different agents in one market.
  • There are no commissions, no borrow fee on a short and no stop orders. A stop you check at each step fills a median 26.5 bp past its level at six steps a day. Uninvested cash earns nothing unless you pass cash_interest=True, and a negative cash balance pays the policy rate, which is below a broker's margin rate.

docs/AGENTS.md has the measurements behind each of these.

Examples

The twelve numbered examples/ are in reading order, and the test suite runs them:

00-a-year-in-one-marketStart here: one company, one year, two crises, one chart
01-first-simulationUniverse, engine, order book, determinism
02-evaluating-a-strategySpecs, baselines, ranking across seeds
03-why-did-the-price-moveThe eleven factors that sum to the mispricing's move
04-how-realistic-is-thisThe realism panel and the limits
05-training-an-agentThe Gymnasium environment, and what size costs
06-execution-and-impactTCA and the counterfactual run
07-research-workflow.pyA whole study in one file. It takes about forty seconds of CPU and needs tradefloor[arrow]
08-claude-agent.pyAn LLM agent trading the market through the harness
09-a-pandemic-shaped-marketA real 2020-21 macro path, and which fields transmit. Pinned to pt-v12, whose QE channel carries the valuation path, with the same path on the default, pt-v20, at the end
10-forking-a-marketFork a market, raise the rate in one branch, and compare the futures
11-scenario-fork.pyA scenario file applied to one branch of a fork, and what it cost

The rate-shock/ study is the demo above, and integrations/ runs the same kind of experiment through each agent framework, offline and without an API key.

Documentation

https://docs.tradefloor.dev covers install, the API, the guides and how the model is measured. These pages in this repository have the detail behind the sections above:

Contributing and support

CONTRIBUTING.md explains how to build and test the project. Its main rule is that any change to the simulated trajectory is a breaking change, however small, so a model change ships as a new preset. RELEASING.md is the release checklist.

Report a vulnerability through GitHub's security advisory form, not a public issue. SECURITY.md says what is in scope. Bugs and questions go to GitHub issues.

0.8.5 and the 0.8 patches after it are the long-term support line, with bug and security fixes for 24 months. docs/SUPPORT.md says what a support line promises and which release to pin for a long study.

Citing tradefloor

Cite the version you ran and name the preset. The same version can run several presets, and results depend on the preset.

bibtex
@software{tradefloor,  author  = {Coombes, Simon},  title   = {tradefloor: a deterministic market simulator with a limit order book},  version = {0.9.0},  year    = {2026},  url     = {https://github.com/simoncoombes/tradefloor},  note    = {Model preset pt-v20}}

CITATION.cff carries the same details, and GitHub's "Cite this repository" button reads it.

In the text, say which model you used, for example: "tradefloor 0.9.0, preset pt-v20, specified in its docs/MODEL.md". docs/REPRODUCIBILITY.md says how to publish a result so a reader can rerun it, and how to show a score was not tuned to its seeds.

License

tradefloor is licensed under MIT OR Apache-2.0, at your option. See LICENSE-MIT and LICENSE-APACHE. GitHub's sidebar reads Apache-2.0 because its license detection picks one file and stops. The grant that applies is the dual one, stated in pyproject.toml, rust/Cargo.toml and this section.

来源:README.md,提交 d09da12

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.9.0最新Oct 3, 2026