TY - GEN
T1 - Learning to Trade with Preferences
T2 - 6th ACM International Conference on AI in Finance, ICAIF 2025
AU - Xu, Haohan
AU - Bohne, Jason
AU - Polak, Pawel
AU - Byrd, David
AU - Rosenberg, David
AU - Kazantsev, Gary
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2025/11/14
Y1 - 2025/11/14
N2 - Deterministic execution strategies - TWAP, VWAP, Implementation Shortfall (IS), and Percent-of-Volume (POV) - remain widely used in institutional trading due to their interpretability, alignment with client benchmarks, and regulatory transparency. However, they are inflexible to changing market microstructure, and brokers often switch or combine them manually in response to prevailing conditions. We propose a reinforcement learning framework that constructs adaptive, interpretable policies as state-dependent mixtures of these deterministic strategies. Using the ABIDES-Gym multi-agent simulator, we first pretrain each strategy independently via tabular -learning to adapt locally to execution feedback. We then apply Mixture-of-Experts Direct Preference Optimization (MoE-DPO) to fine-tune and integrate these policies through preference-based optimization over execution trajectories. This enables high-level customization of execution policies based on broker preferences that could be difficult to encode via explicit reward functions. Empirical evaluations demonstrate that MoE-DPO consistently outperforms both standalone -learned strategies and their adaptive mixtures in terms of realized profit and execution cost across diverse market regimes. These results highlight the potential of preference-based learning for aligning execution algorithms with human supervisory signals in dynamic trading environments.
AB - Deterministic execution strategies - TWAP, VWAP, Implementation Shortfall (IS), and Percent-of-Volume (POV) - remain widely used in institutional trading due to their interpretability, alignment with client benchmarks, and regulatory transparency. However, they are inflexible to changing market microstructure, and brokers often switch or combine them manually in response to prevailing conditions. We propose a reinforcement learning framework that constructs adaptive, interpretable policies as state-dependent mixtures of these deterministic strategies. Using the ABIDES-Gym multi-agent simulator, we first pretrain each strategy independently via tabular -learning to adapt locally to execution feedback. We then apply Mixture-of-Experts Direct Preference Optimization (MoE-DPO) to fine-tune and integrate these policies through preference-based optimization over execution trajectories. This enables high-level customization of execution policies based on broker preferences that could be difficult to encode via explicit reward functions. Empirical evaluations demonstrate that MoE-DPO consistently outperforms both standalone -learned strategies and their adaptive mixtures in terms of realized profit and execution cost across diverse market regimes. These results highlight the potential of preference-based learning for aligning execution algorithms with human supervisory signals in dynamic trading environments.
KW - ABIDES
KW - Algorithmic Trading
KW - Direct Preference Optimization
KW - Mixture-of-Experts
KW - Order Execution
UR - https://www.scopus.com/pages/publications/105023064278
U2 - 10.1145/3768292.3770390
DO - 10.1145/3768292.3770390
M3 - Conference contribution
AN - SCOPUS:105023064278
T3 - ICAIF 2025 - 6th ACM International Conference on AI in Finance
SP - 762
EP - 770
BT - ICAIF 2025 - 6th ACM International Conference on AI in Finance
PB - Association for Computing Machinery, Inc
Y2 - 15 November 2025 through 18 November 2025
ER -