DeepSeek Dominates Stock Trading Test, But ChatGPT Rules Event Prediction

DeepSeek beats AI rivals in stock trading, delivering 73% returns. Discover the winner and what it means for investors.
Robot with blue visor on stock exchange floor, financial screens, Citadel Securities logos, symbolizing DeepSeek's test.
Robot with blue visor represents DeepSeek's stock test. By Andres SEO Expert.

Key Takeaways

  • DeepSeek portfolios returned 73% since launch in 2023, outperforming ChatGPT and Grok year-to-date.
  • Independent research from AIMultiple shows ChatGPT 5 Thinking leads in financial event prediction, highlighting task-specific model strengths.
  • The findings underscore that no single AI model excels universally; selection depends on the specific financial task.

AI Trading Showdown: Professor’s Three-Year Test Crowns DeepSeek

Dr. Alejandro Lopez Lira, an associate professor of finance at the University of Florida, has spent nearly three years testing whether AI chatbots can deliver market-beating stock returns. His conclusion after evaluating DeepSeek, ChatGPT, Grok, and Claude: DeepSeek is the clear winner for 2026. ‘DeepSeek is by far the best trader,’ Lopez Lira said in an interview with Business Insider.

Lopez Lira’s portfolios are live on Autopilot, a copy-trading platform that allows retail investors to mirror the trades of successful investors. Chris Josephs, co-founder of Autopilot, noted increasing interest in AI-managed portfolios. The test results show DeepSeek’s portfolio up 35% year-to-date and 73% since its launch in 2023. ChatGPT returned 78% since inception but only 16% in 2026, while Grok gained 54% since launch and 19% year-to-date.

DeepSeek’s Performance Edge: By the Numbers

DeepSeek’s portfolio currently holds Micron Technology as its top position, a bet shared by Grok. However, DeepSeek recently increased its stake in Nvidia while reducing exposure to Vista Energy. ChatGPT’s portfolio, by contrast, leads with Kratos Defense & Security Solutions, a lesser-known aerospace firm recently cited by Stifel as a top pick. Claude’s portfolio takes a more conservative approach, with its largest holding in the iShares 0-3 Month Treasury Bond ETF and a significant position in Vista Energy — a contrarian stance given DeepSeek’s reduction.

According to the Business Insider report, Lopez Lira estimates that the AI models make the same trading decision only about 50% of the time, despite processing similar information. ‘There’s some setup involved — you have to add the latest information — but the final decision is the model’s completely,’ he explained. His white paper on Autopilot includes examples of the prompts used to steer each model.

Real-Time Research Validates the Complexity of AI Trading Models

Independent research by Ezgi Arslan, PhD, at AIMultiple tested 14 generative AI models on a different financial task: predicting stock market reactions to death events in family firms. The study required models to label expected returns as ‘significantly positive’, ‘significantly negative’, or ‘not significant’ based on firm financials and founder details.

Results show that ChatGPT 5 Thinking achieved the highest accuracy at 74%, followed by Gemini 2.5 Pro at 71%. DeepSeek V3.2 scored 58% in the base round, improving to 64% when given additional contextual data. Meanwhile, Claude Sonnets scored between 46–48%, and some earlier GPT and Gemini models performed worse. Notably, adding extra information reduced accuracy for several models, indicating that more data can sometimes confuse simpler architectures.

These findings complement Lopez Lira’s real-world stock picking results. While DeepSeek excels in active trading, ChatGPT variants demonstrate superior performance in structured classification tasks. This suggests that the ‘best’ AI model depends heavily on the specific use case — a critical insight for investors and developers building financial AI systems.

Implications for Investors and the Future of AI in Finance

Lopez Lira’s experiment and the AIMultiple study together paint a nuanced picture: AI models are not interchangeable. The same model that beats the market in one context may lag in another. For investors, this means that relying on a single AI for trading decisions is risky. A diversified approach — leveraging multiple models for different tasks — may yield more consistent results.

As AI capabilities evolve rapidly, today’s champion could be overtaken. The finance professor’s methodology, which combines human-supplied data with model autonomy, highlights the importance of human oversight. For developers and fintech firms, the takeaway is clear: rigorous, ongoing benchmarking across diverse tasks is essential to match the right model with the right job.

For those building AI-powered applications or seeking to optimize their digital presence, expert guidance is essential. Andres SEO Expert’s programmatic SEO and AI automation services can help businesses stay ahead. Connect with Andres to explore how these technologies can be leveraged for your growth. Learn more about Andres SEO Expert‘s approach.

Frequently Asked Questions

What is the key finding of Professor Lopez Lira’s three-year test of AI trading chatbots?

DeepSeek is the best performer among AI models for stock trading, with its portfolio up 35% year-to-date and 73% since launch in 2023.

How does DeepSeek’s performance compare with ChatGPT, Grok, and Claude in stock trading?

DeepSeek outperformed: ChatGPT returned 78% since inception but only 16% year-to-date; Grok gained 54% since launch and 19% year-to-date; Claude took a conservative approach with bond holdings.

What methodology did Lopez Lira use to test the AI models?

He provided each model with the same latest information and prompts, then let the model make its own trading decisions without human intervention.

What did the AIMultiple study reveal about AI models in financial prediction tasks?

ChatGPT 5 Thinking achieved highest accuracy (74%) for predicting stock reactions to death events, while DeepSeek scored 58-64%. Adding more data sometimes reduced accuracy for simpler models.

What are the implications for investors using AI for trading?

AI models are not interchangeable; the best model depends on the specific task. A diversified approach using multiple models for different tasks may yield more consistent results.

How can businesses leverage AI for financial decision-making?

Businesses should conduct rigorous benchmarking across diverse tasks and maintain human oversight. Expert guidance from services like programmatic SEO and AI automation can help optimize AI use.

Should investors rely on a single AI model for trading?

No, relying on a single AI is risky. A diversified approach leveraging multiple models for different tasks may yield more consistent results, given that even top models vary by task.

Prev Next

Subscribe to My Newsletter

Subscribe to my email newsletter to get the latest posts delivered right to your email. Pure inspiration, zero spam.
You agree to the Terms of Use and Privacy Policy