Skip to main navigation Skip to search Skip to main content

Learning to Trade with Preferences: Interpretable Execution via Mixture-of-Experts

  • Haohan Xu
  • , Jason Bohne
  • , Pawel Polak
  • , David Byrd
  • , David Rosenberg
  • , Gary Kazantsev
  • Stony Brook University
  • Bowdoin College
  • Bloomberg L.P.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Deterministic execution strategies - TWAP, VWAP, Implementation Shortfall (IS), and Percent-of-Volume (POV) - remain widely used in institutional trading due to their interpretability, alignment with client benchmarks, and regulatory transparency. However, they are inflexible to changing market microstructure, and brokers often switch or combine them manually in response to prevailing conditions. We propose a reinforcement learning framework that constructs adaptive, interpretable policies as state-dependent mixtures of these deterministic strategies. Using the ABIDES-Gym multi-agent simulator, we first pretrain each strategy independently via tabular -learning to adapt locally to execution feedback. We then apply Mixture-of-Experts Direct Preference Optimization (MoE-DPO) to fine-tune and integrate these policies through preference-based optimization over execution trajectories. This enables high-level customization of execution policies based on broker preferences that could be difficult to encode via explicit reward functions. Empirical evaluations demonstrate that MoE-DPO consistently outperforms both standalone -learned strategies and their adaptive mixtures in terms of realized profit and execution cost across diverse market regimes. These results highlight the potential of preference-based learning for aligning execution algorithms with human supervisory signals in dynamic trading environments.

Original languageEnglish
Title of host publicationICAIF 2025 - 6th ACM International Conference on AI in Finance
PublisherAssociation for Computing Machinery, Inc
Pages762-770
Number of pages9
ISBN (Electronic)9798400722202
DOIs
StatePublished - Nov 14 2025
Event6th ACM International Conference on AI in Finance, ICAIF 2025 - Singapore, Singapore
Duration: Nov 15 2025Nov 18 2025

Publication series

NameICAIF 2025 - 6th ACM International Conference on AI in Finance

Conference

Conference6th ACM International Conference on AI in Finance, ICAIF 2025
Country/TerritorySingapore
CitySingapore
Period11/15/2511/18/25

Keywords

  • ABIDES
  • Algorithmic Trading
  • Direct Preference Optimization
  • Mixture-of-Experts
  • Order Execution

Fingerprint

Dive into the research topics of 'Learning to Trade with Preferences: Interpretable Execution via Mixture-of-Experts'. Together they form a unique fingerprint.

Cite this