Industry · · 13 min read

Will prediction markets turn your platform into a bot target?

Will prediction markets turn your platform into a bot target?

Prediction markets can give bots a direct financial incentive to manipulate online metrics.

TL;DR

Prediction markets such as Polymarket and Kalshi let people bet on future events, from election results to sports outcomes.

Some markets are resolved using metrics published by websites and applications: Spotify streams, Twitch views, public votes, player counts, or crowdsourced leaderboard rankings.

These metrics come from user activity, so bots can often influence them too.

A trader may profit from inflating a song, a streamer, a game, or an AI model without having any connection to the platform or the person receiving the artificial activity. The money comes from the prediction market.

This gives platforms a new type of abuse to consider. A public ranking or counter can become the settlement source for a financial market without the platform agreeing to it or even knowing that the market exists.

Prediction markets give bots new targets

Prediction markets have grown quickly over the last few years. Platforms such as Polymarket and Kalshi let users buy positions on the outcome of future events.

Many markets cover elections, sports, economic decisions, company announcements, and other events that online users cannot easily change.

Others use metrics published by online platforms to decide the winner. A market may resolve based on which song ranks first on Spotify, which streamer records the most viewing hours on Twitch, or which AI company occupies the top position on a public leaderboard.

A Spotify stream, a Twitch viewer, or a leaderboard vote can therefore affect the payout of a financial market. Someone who holds the right position may have a reason to generate that activity artificially.

The trader’s edge comes from influencing the event itself rather than learning its outcome early. A person can place a bet on a song reaching number one, then try to generate enough streams to push it there.

I started looking into this after reading about suspicious Spotify streams around a Kalshi market. The case involved a song that briefly reached the top of a US Spotify chart before Spotify removed more than 500,000 streams it considered artificial.

The timing raised doubts about whether the streams were linked to the bet. There is no public proof that they were, but the mechanism was credible enough to make me look for other examples.

I found markets based on Twitch viewing hours and Arena AI leaderboard positions. Similar markets could use Reddit votes, Roblox player counts, GitHub stars, application rankings, or public polls.

All of these platforms already deal with bots. Prediction markets can make certain metrics more valuable to manipulate, especially when the market is large, the outcome is close, and the artificial activity only needs to survive until a specific resolution time.

The platform receives the fake traffic, distorted analytics, and investigation work. The trader collects the profit somewhere else.

The Spotify case

In June 2026, “Earrings,” a song by Malcolm Todd, experienced a sudden increase in US Spotify streams. The increase briefly pushed it to the top of Spotify’s daily US chart.

At the same time, Kalshi offered a market on whether Malcolm Todd would have a number-one song on Spotify USA before the end of June. Traders had previously priced that outcome at roughly 2.5 percent.

Spotify later removed more than 500,000 artificial streams and corrected the song’s chart position from first to fourth. Kalshi had already resolved the market using the initial chart. Traders who bought the unlikely outcome could have made around 20 times their initial wager. There is no evidence that Malcolm Todd or his team were involved.

Spotify defines an artificial stream as one that does not reflect genuine listening intent, including streams generated through bots or scripts. These streams have traditionally been used to collect royalties, sell promotion services, or influence recommendation systems.

A prediction market offers another source of revenue. The person buying the streams does not need to own the song or know the artist. They only need a market position that becomes more valuable when the song moves up the chart.

The case also shows how timing can favor the attacker. Spotify detected and corrected the artificial activity, but the correction arrived after the market had settled.

For the attacker, permanent manipulation may be unnecessary. The fake activity only needs to remain visible until the resolution snapshot.

Manipulation has to be profitable

The possibility of influencing a metric does not automatically make it an attractive target.

The operation needs positive expected value.

The attacker has to pay for automation, accounts, proxies, devices, CAPTCHA solving, verification, monitoring, and the market position itself. They also face the risk that the platform filters the activity, the market disputes the outcome, or the metric moves naturally in the opposite direction.

The economics depend on three main factors:

  1. How much the operation can increase the probability of a favorable result
  2. How much profit is available from the market
  3. How much the operation costs, including the risk of detection and failure

A campaign that increases the probability of success from 20 percent to 70 percent creates value even though it does not guarantee the outcome. The relevant gain comes from that 50 percentage point improvement.

Close rankings and narrow thresholds are particularly attractive.

If two streamers are separated by a small number of viewing hours, the attacker only needs to manufacture the difference. If a song is close to entering the top ten, the attacker can wait for legitimate listeners to generate most of the required streams.

The cost also depends on whether the attacker can reuse the infrastructure.

A one-off operation may require building automation, acquiring accounts, and learning how the platform’s defenses behave. Those costs fall when the same platform is repeatedly used as a resolution source.

An attacker could maintain a pool of aged accounts on several popular platforms and activate them when a suitable market appears. The accounts might remain mostly dormant between operations, then be used for streams, votes, views, or other interactions around a resolution deadline.

This turns manipulation into a repeatable business rather than a single campaign. Account creation, proxy integration, device management, and platform research can be amortized across several bets.

Platforms that frequently appear in prediction-market rules may therefore become more attractive long-term targets. Attackers have a reason to industrialize account creation and maintain access before a profitable opportunity appears.

Some metrics are easier to target than others

The most attractive metrics tend to share a few properties.

Users directly generate them. A stream exists because someone plays a song. Viewing hours increase when someone watches a live stream. A leaderboard vote exists because a user chooses one response over another.

The action can also be repeated or distributed. An attacker may generate it across accounts, devices, sessions, or IP addresses.

The market structure matters too. A point-in-time ranking is often easier to target than a large annual total. The attacker can concentrate resources around a known deadline and stop once the market has resolved.

Detection and correction delays create another opportunity. Many platforms publish counters quickly, then investigate suspicious activity later. The public value at noon may differ from the corrected value two days later.

Finally, the market needs enough liquidity to finance the campaign. A sophisticated operation will not make sense when the maximum realistic profit is a few hundred dollars.

The riskiest combination is a large market, a close outcome, a user-generated metric, and a precise resolution time.

AI leaderboards as a resolution source

One Polymarket market asks which company will have the best AI model at the end of 2026.

Its rules specify that the market will resolve using the company whose model holds the highest position on the Arena text leaderboard at 12:00 PM Eastern Time on December 31, 2026.

Arena ranks AI models through crowdsourced pairwise comparisons. A user submits a prompt and receives responses from two anonymous models. The user votes for the response they prefer, and Arena reveals the model names afterward. These votes contribute to the public leaderboard.

The models being anonymous makes targeted voting more difficult, but their outputs are not always indistinguishable.

I tested this by repeatedly asking prompts with a similar structure, including variations of:

What is the best way to win on Polymarket?

After each vote, Arena revealed which models had produced the answers. Some company-level patterns quickly became visible. Gemini responses often followed a similar structure and tone, while Claude responses looked noticeably different. In particular, we quickly notice that Google starts its answer with “Winning” or “Winning on Polymarket” whereas Antropic’s answer structure looks way different.

This was a small, informal test, so it does not show that models can be identified reliably from every answer. The same model can respond differently across prompts, and models from different companies can produce similar outputs.

An attacker would not need perfect identification, though.

They could repeat prompts with similar structures, observe the model names after each vote, and build a collection of labeled examples. Over time, they could learn which formatting choices, refusal styles, phrases, or answer structures are more commonly associated with a company.

They could then vote only when sufficiently confident and skip ambiguous comparisons. Even a weak ability to identify the target model above random chance could become useful when repeated across enough votes.

The post-vote reveal makes this learning process easier. Every comparison gives the participant another labeled example they can use to refine future guesses.

This still adds cost and uncertainty. The attacker needs more queries, more accounts, and enough successful identifications to affect the ranking. Arena also has defenses against coordinated or automated voting.

Researchers have already explored how coordinated voting could influence Arena-style leaderboards. One paper analyzed 1.7 million historical Chatbot Arena votes and showed that, in offline simulations, some vote-rigging strategies could improve a model's ranking with only a few hundred additional votes (Improving Your Model Ranking on Chatbot Arena by Vote Rigging). Another study demonstrated that an insufficiently protected simulated leaderboard could be manipulated with roughly a thousand votes, then worked with the Arena team on mitigations such as rate limiting, login requirements, bot protection, and malicious-user detection (Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards).

These studies do not tell us how much it would cost to manipulate Arena today. Its defenses change the economics, and my small experiment does not measure whether a real campaign would succeed.

They do show that crowdsourced leaderboards have an adversarial voting problem. Prediction markets can make that problem more valuable by attaching a direct financial payout to a company’s position at a specific time.

The trader funding the activity would not need to work for Google, OpenAI, Anthropic, or any other AI company. They would only need a prediction-market position that becomes more valuable when one company moves up or down the leaderboard.

Twitch viewing hours

Another Polymarket market asks which Twitch streamer will be the most watched during July 2026.

It resolves using the channel ranked first by total hours watched on TwitchTracker’s 30-day leaderboard at 12:00 PM Eastern Time on July 31.

Hours watched comes directly from viewer activity. A user who remains connected to a stream contributes to the measured total.

Viewbotting is already a known form of abuse on Twitch. Twitch says it investigates artificially inflated viewers, chat activity, and follower counts, and collects information about users operating these services.

Prediction markets expand the group of people who may pay for viewbotting.

The streamer may want higher visibility, more social proof, or better sponsorship metrics. A trader may only want one channel to finish above another before the resolution deadline.

The streamer may have no involvement at all. The channel receiving the artificial viewers is not necessarily the attacker or even a willing participant.

This complicates investigations. The platform sees viewbotting, but the financial motive exists outside the platform and may have no clear link to the account receiving the traffic.

Reddit, Roblox, and other potential targets

Spotify, Twitch, and Arena provide concrete examples, but the same incentive can apply to many platforms.

Platform Metric that could resolve a market Possible automated influence
Reddit Post votes, subreddit members, comments, ranking position Account farms, coordinated voting, automated posting
Roblox Concurrent players, visits, favorites, likes Automated clients, multi-accounting, idle sessions
GitHub Repository stars, forks, issue activity Fake, purchased, or compromised accounts
YouTube Views, subscribers, likes, watch time Automated viewing and account farms
Discord Server members, reactions, poll results Account farms and scripted interactions
App stores Downloads, ratings, category position Install farms and coordinated reviews
Online games Concurrent players, match counts, rankings Automated game clients and multi-accounting
Public polls Votes and candidate rankings Automated voting and fake accounts

Not every metric is equally vulnerable.

Reddit filters vote manipulation and explicitly includes interference with upvote and downvote totals in its content-manipulation category. In July 2026, Reddit also said its systems were revoking nearly two million fake votes each day.

Roblox and online games may require more expensive automation because the attacker needs to run real clients, consume computing resources, and avoid anti-cheat systems.

App stores can require devices, established accounts, and payment methods. GitHub stars may be cheap to generate with new accounts, but platforms can later remove suspicious activity. These controls change the cost but rarely reduce it to infinity.

A market can still become attractive when the payout is large enough, the gap is small enough, or the attacker already owns the required infrastructure.

Staying under the radar

The most effective campaign may not produce a large spike.

An attacker trying to win a prediction market only needs to move the metric far enough to change the result.

If two models have similar leaderboard scores, a small shift may be enough. If two streamers are close near the end of the month, the attacker only needs to cover the remaining gap.

The operation can spread activity across accounts, devices, networks, and time. Each account may generate a small amount of activity that looks reasonable in isolation.

A reusable pool of accounts makes this easier. Accounts can be aged gradually, given normal-looking histories, and kept ready for future markets. Instead of creating thousands of accounts immediately before every bet, the attacker develops an inventory that can be instrumented when needed.

The known resolution deadline gives the attacker another advantage. They know the exact threshold, the remaining gap, and how long the artificial activity needs to survive.

The platform may only see a modest change in otherwise normal-looking traffic.

Your platform becomes an involuntary referee

Prediction markets need an external source to decide which positions win.

For markets tied to online metrics, that source may be a public chart, leaderboard, analytics page, or third-party tracker built from platform data.

Spotify’s chart, Arena’s leaderboard, or TwitchTracker’s ranking becomes financially authoritative because the market rules say so.

The platform may not know that the market exists. It did not agree to provide a settlement service, and its abuse-detection process may not match the market’s deadline.

It still absorbs the impact.

Artificial activity pollutes growth, engagement, retention, and conversion data. Product teams may use those numbers before security teams identify the problem.

The activity can also influence recommendations. A song, stream, post, game, or application pushed into a discovery surface may attract real users. Organic engagement then mixes with the original artificial traffic.

Streams, watch time, active players, and votes may affect royalties, advertising, creator programs, rankings, or internal promotion. A trader’s attempt to win an external bet can change how the platform distributes money and attention.

The platform also pays the infrastructure and moderation costs. It receives the requests, accounts, investigations, appeals, and support tickets.

The prediction market receives the trading fees.

The market may settle before the platform corrects the metric

Platforms often need time to identify coordinated manipulation.

Suspicious activity may become visible only after several accounts are linked, a proxy network is identified, or a longer behavioral pattern emerges. Some cases require manual investigation.

Prediction markets usually define a specific resolution time. The rules may say that a page will be checked at noon on a certain date and that the displayed value will determine the winner.

The real-time value can differ from the value after automated filtering or a fraud investigation.

The Spotify case followed this pattern. The song initially appeared in first place, Kalshi settled the market, and Spotify later corrected the chart to place it fourth.

A fixed snapshot gives the attacker a clear operational goal: keep the manipulated activity visible until the check takes place.

The activity can be detected and removed later, after the payout has already been made.

What platforms should monitor

Prediction markets don't require a completely new detection strategy.

The same techniques used to detect fake accounts, engagement fraud, and coordinated abuse remain relevant. Large platforms should continue monitoring for:

The difference is that campaigns may become harder to explain.

Historically, suspicious activity often had an obvious objective. Fake accounts abused promotions. Streaming bots generated royalties. Viewbots inflated popularity. Engagement bots sold followers, likes, or comments.

Now, some campaigns may have no obvious benefit on the platform itself.

A few thousand fake streams, votes, or viewers may not generate meaningful revenue, improve a creator's reputation, or noticeably change user growth metrics. Their only purpose may be to move a public metric just enough to influence the outcome of a prediction market.

That makes correlation even more important.

Looking at individual accounts is rarely enough to understand a sophisticated operation. The interesting signals usually emerge when activity is analyzed across the entire attack infrastructure: accounts created with similar email patterns, related device fingerprints, shared proxy networks, synchronized behavior, or infrastructure that quietly persists for months before becoming active around a particular event.

This also changes the economics for attackers.

If a platform is frequently used as the resolution source for prediction markets, building fake accounts becomes a long-term investment instead of a one-off operation. Rather than creating new accounts for every campaign, attackers can maintain pools of aged accounts across multiple platforms and activate them whenever a profitable opportunity appears.

The underlying detection problem remains the same: identify coordinated abuse before it affects metrics that users and other systems trust. Prediction markets simply increase the value of manipulating those metrics.

Prediction markets need to price manipulation risk

Prediction-market platforms also have a role.

A public counter should not automatically be treated as an objective and final fact. Its reliability depends on how it is generated, how expensive it is to manipulate, and how quickly the source platform corrects fraud.

Useful safeguards include:

The objective is to push the expected return below the cost and risk of the bot operation.

A new incentive for an old bot ecosystem

Streaming bots, viewbots, fake accounts, coordinated votes, install farms, and automated game clients already exist.

Prediction markets give that infrastructure more potential customers, targets, and revenue.

A platform may discover that one of its public metrics now controls the payout of a large external market. Attackers may prepare accounts months in advance, reuse the same automation across several bets, and keep each campaign small enough to avoid obvious detection.

The platform may see more automated activity without understanding the motive. It may be expected to preserve the integrity of a financial product it never agreed to support.

Once your platform’s metric becomes the answer to a bet, your platform becomes part of the attack surface.

Read next