German Bank Lets Claude and ChatGPT Trade for Customers; Claude Beat Human Traders 76% of the Time

Published August 26, 2026By ABD Legacy LLC
AI in finance Claude trading performance AI agents in finance ChatGPT trading regulated AI deployment
German bank Scalable Capital lets ChatGPT, Claude and Grok trade for customers in AI in finance 2026

On August 25, 2026, Scalable Capital, a Munich-based digital bank with more than €60 billion in client assets, became the first bank in Europe to open its platform to major AI assistants. Customers can now connect ChatGPT, Claude, or Grok to their brokerage account and let them place trades and set up savings plans in plain language. [2][4]

The same news cycle carried a bolder claim: Claude beat human traders 76% of the time. That number is real, but it did not come from the bank's test. It comes from a separate simulated study by Elm Wealth. Two events, one question: how much should anyone trust AI with real money? [2][3]

For anyone selling AI in finance 2026, this launch and this study show what agentic AI can do, and how far it is from managing risk on its own.

A Major German Bank Just Put ChatGPT and Claude in the Trading Workflow

Scalable Capital is not a small fintech experiment. Founded in 2014, it manages more than €60 billion in client assets for over a million customers. On August 25, 2026, it launched "Agentic Investing," claiming a first: no other bank in Europe had opened its platform to major AI assistants. [2][4] TechTimes put it more sharply: Scalable is Europe's first ECB-licensed bank to turn a general-purpose AI chatbot into a live trade executor on a regulated retail brokerage account. [5]

In natural language, customers can execute trades, set up savings plans, manage watchlists, and search stocks and ETFs. Scalable Insights analytics, including diversification checks and risk assessments, are exposed to agents too. [4]

The launch landed alongside a Fortune headline claiming Claude beat human traders 76% of the time. The catch: that number belongs to a different firm's simulation.

The Test: A Crystal Ball With the Prices Blacked Out

The 76% figure comes from Elm Wealth's "Crystal Ball Challenge," published July 1, 2026. Fortune's coverage and the bank launch are two separate events. Scalable did not run this test. [2][3]

Elm staked 120 finance-trained adults $50 each. Each saw the Wall Street Journal's front page a day before publication, with market moves blacked out, and could go long or short the S&P 500 and 30-year Treasuries with leverage. Each player got 15 opportunities, one front page per year from 2008 to 2022. [3]

Four AI models played the same game: Claude, ChatGPT, Gemini, and Grok, each starting with $1 million in play money and running 10 rounds. They saw the same headlines, transcribed by OCR, reasoned from first principles without knowing actual returns, and were monitored for compliance (except Grok). Across roughly 200 sessions, the results were striking. [3]

The whole exercise was simulated, and Elm disclaimed it as entertainment and education, not investment advice. [3]

Capability Numbers: Claude 76%, ChatGPT 63%, and Where They Fall Short

Can AI beat human traders? In one controlled test, yes. Claude beat human players in 76% of sessions and ChatGPT in 63%. But the same study found every model took too much risk on position sizing. Beating humans on direction is not the same as protecting capital.

Claude beat human players in 76% of sessions, ChatGPT in 63%, Gemini in 43%, and Grok in 51%. [2]

Win rate vs human traders, Elm Wealth Crystal Ball Challenge Claude 76% ChatGPT 63% Grok 51% Gemini 43% Simulated sessions, play money — not investment performance.

Claude beat human players in 76% of sessions in Elm Wealth's Crystal Ball Challenge; ChatGPT 63%, Grok 51%, Gemini 43%.

The more useful finding: what the AIs were good at and what they were not. Elm concluded they were "roughly as good as expert humans on the what decision": connecting macro news to near-term market direction. A reference group of five professional macro traders averaged $2.3 million in ending wealth. Claude's Sharpe ratio, roughly 0.3, was the highest; Gemini and Grok had positive ratios yet lost money on average, from over-betting rather than bad calls. [3]

Strong on what to invest in, weak on how much to invest. Every model knew the Kelly criterion and the Merton share in the abstract; none applied it to position sizing. High-reasoning Claude reduced its position sizes, while high-reasoning Gemini and ChatGPT took more risk than Elm called appropriate. For anyone tracking Claude trading performance, this is the most credible public benchmark yet. [3]

How the Bank Actually Lets an AI Trade: MCP, Approvals, and the Fine Print

The mechanics matter, because "ChatGPT trading" in a real product looks nothing like the demo. Scalable offers two routes: an MCP server for cloud AI assistants and a CLI for locally hosted models. MCP is the open standard for connecting AI assistants to external tools. [4][5][7]

The flow: the customer asks in natural language, the AI calls a standardized tool mapped to Scalable's brokerage API, Scalable returns a confirmation request, and the user confirms. Only then does the trade execute. The MCP server runs on Scalable's infrastructure, the model on OpenAI's, Anthropic's, or xAI's, talking over JSON-RPC 2.0. [5]

The fine print does the real reassuring. Users must approve trades and savings plans before execution. Agents cannot make payments or withdraw money. Two-factor authentication is required, Key Information Documents are served before every transaction, and Agentic Investing can be deactivated anytime. [2][4][7]

Scalable's chief product officer, Alexander Siepp, stressed this is not a formal partnership with OpenAI or Anthropic; MCP is an open technology, and Scalable uses "ready-made technology integrations offered by those AI assistants." OpenAI, Anthropic, Google, and xAI did not respond to Fortune's request for comment. [2] The fine print also says the CLI and MCP are "technical interfaces to external AI applications" outside Scalable's control, and AI outputs are executed "at the investor's own risk." [4]

Scalable is not alone: Robinhood launched Agentic Trading in May 2026 with about 100,000 accounts by the end of Q2, and eToro's Tori executed more than 500,000 trades in its first year. AI agents in finance are moving from demos to deployed products, the same agentic commerce shift agencies have been tracking. [5]

The Regulatory Catch: MiFID II, ESMA, and Who Is Liable

Regulated AI deployment is where the demo meets its limits. Under MiFID II, algorithmic trading is trading where a computer algorithm automatically determines individual parameters of orders. ESMA's February 2026 briefing clarified this applies even where a human intervenes at other stages, told firms using AI in trading to reflect it in self-assessments and align with the EU AI Act. Non-binding, but a clear statement of expectations. [5]

Scalable's architecture assigns the decision-maker role to the human: the AI is the execution layer, and the customer approves each order. Whether regulators accept that framing as agents grow more autonomous is an open question. [5]

There is also a non-market security catch. Prompt injection and tool poisoning are documented vulnerability classes for MCP-connected agents; the confirmation step and 2FA reduce but do not eliminate the risk, and the NSA published MCP security guidance in May 2026. [5] For US businesses, state AI law is moving: our AI compliance explainer covers SB 53, and the prompt injection attack class should worry anyone connecting an agent to money.

What This Means for Agency Clients

The client conversation has changed. "Your AI can now touch money" is a product feature at a regulated European bank. The Elm risk data is the counterweight to the hype: AIs sized stock positions at 7x to 12x on average, with daily volatility of 20% to 40%, a level Elm called too much risk of catastrophic loss. [2][3]

Every agency client deploying an AI agent that can act, not just answer, should be asked four questions:

  1. Who approves what? Which actions require human confirmation, and is that confirmation real? An AI agent permissions audit answers this.
  2. Can the agent move money out? In Scalable's design, no. In less careful integrations, the answer is not always no.
  3. What does the fine print say about liability? Scalable explicitly disclaims responsibility for AI-generated outputs. Most vendors will.
  4. What happens when a prompt-injected tool call arrives? This is the core AI agent security risk in MCP-connected systems, and it deserves a review before launch.

These are the same failure modes agencies handle in automation work, from runaway costs to identity checks. In finance they carry a bigger price tag: a data exposure now has a dollar sign. [2][3][4]

The Risk-Focused Conclusion: AIs Can Call Trades, Not Manage Risk

Put the two events side by side and the pattern is clear. AI models can pick a direction roughly as well as expert humans. They cannot yet size a position responsibly. Capability is not capital preservation. [2][3]

Elm's math makes the risk concrete. Since 2000, the market has moved more than 5% on 23 days and more than 9% on seven days; at 7x to 12x sizing, a bad call meeting a big move meant catastrophic loss. Grok would have ended with $1.27 million if it cut bets 60%; Gemini $1.005 million with 90% smaller bets. [3]

Even Siepp, whose company sells the feature, hedged: whether customers "end up finding the holy grail together with your AI assistant on high returns and low risks or not," he told Fortune, "remains to be seen." [2]

The takeaway for agencies is governance, not magic: approvals, audits, and a kill switch. The AI agents trading guardrails we documented when Binance launched Agent OS apply here with more money on the line. Start with an AI readiness audit and treat position sizing, liability, and prompt injection as design requirements. [3][5]

The demo capability is real. Claude can beat human traders 76% of the time in a simulation. Regulated deployment is a different sport, and the teams that win it respect the difference.

Ready to build the trust layer for agentic finance?

Browse AI Agencies →

Frequently Asked Questions

Can AI agents trade stocks?

Yes. Since August 25, 2026, customers of German bank Scalable Capital can connect ChatGPT, Claude, or Grok over MCP and place trades in natural language. Every trade requires explicit user approval, and agents cannot make payments or withdraw money. [2][4][5]

Did Claude beat human traders?

In the Elm Wealth Crystal Ball Challenge, a simulated test of about 200 sessions, Claude beat human players in 76% of sessions and ChatGPT in 63%. The study also found all four AIs took too much risk on position sizing. It used play money, not real capital. [2][3]

Is AI trading regulated?

Yes. Under MiFID II, algorithmic trading covers any computer algorithm that automatically determines order parameters, even with human involvement. ESMA's February 2026 briefing told firms using AI in trading to reflect it in self-assessments and align with the EU AI Act. [5]

Can ChatGPT execute trades?

Yes, at Scalable Capital since August 25, 2026. Customers connect ChatGPT via MCP to execute trades, set up savings plans, and manage watchlists, with mandatory approval before each order and no ability to withdraw money. [2][4][5]

Is AI trading safe?

Not automatically. Elm Wealth found AIs sized positions at 7x to 12x, risking catastrophic loss. Banks like Scalable add approvals, 2FA, and no-withdrawal limits, but prompt injection remains a documented risk for MCP-connected agents. [3][5]

Sources