SOTA:用期权隐含收益分布引导的期权交易智能体
SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions
用 Qwen 微调出的期权交易智能体,把上千份合约简化成策略选择,六个月样本外赚 18.3%,新闻用法还挺反直觉。
SOTA 是一个期权交易智能体框架,把数千份期权合约抽象为策略级决策,由确定性解析器完成组合落地。该框架基于 Qwen3.8-27B 做监督微调加强化学习后训练。在九只美股大盘股和 SPY 的期权上,六个月样本外测试取得 18.3% 总收益、1.60 夏普比率、8.96% 最大回撤,优于规则型与机器学习策略选择器。研究还发现新闻在监督微调阶段有帮助,但在强化学习阶段保留新闻会让样本外收益从 18.3% 降到 -2.7%。
SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions
As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the agent must decide both which contracts to trade and how to combine them. Existing approaches often sidestep this complexity by restricting the policy to a fixed strategy structure, such as a straddle, limiting their ability to switch strategies as market conditions change. We present SOTA (Stock Options Trading Agents), an agentic trading framework for structured option-strategy selection. SOTA abstracts the large option universe into strategy-level decisions while deterministic resolvers handle portfolio implementation. We develop SOTA by post-training Qwen3.8-27B with supervised fine-tuning followed by reinforcement learning. SOTA is evaluated on options on nine large-cap U.S. equities and SPY against rule-based and machine-learning strategy selectors in the same trading environment. Over a six-month out-of-sample period, SOTA earns an 18.3% total return with a Sharpe ratio of 1.60 and a maximum drawdown of 8.96%. We also document an asymmetric role of news: news improves frontier-teacher trajectories, but retaining news during reinforcement learning reduces out-of-sample return from 18.3% to -2.7%.