Cristiano Ronaldo during Portugal’s losing game against Spain earlier this month. Source: Imagn Images via Reuters Connect
Football fans wagered more than US $14 billion on the FIFA World Cup through prediction markets Polymarket and Kalshi, a Bellingcat analysis has found.
Support Bellingcat
Your donations directly contribute to our ability to publish groundbreaking investigations and uncover wrongdoing around the world.
Donate
On the crypto-based Polymarket, which provid
Cristiano Ronaldo during Portugal’s losing game against Spain earlier this month. Source: Imagn Images via Reuters Connect
Football fans wagered more than US $14 billion on the FIFA World Cup through prediction markets Polymarket and Kalshi, a Bellingcat analysis has found.
Support Bellingcat
Your donations directly contribute to our ability to publish groundbreaking investigations and uncover wrongdoing around the world.
On the crypto-based Polymarket, which provides more information about individual trading accounts than its rival American site Kalshi, we also found that just 1% of users collected the vast majority of winnings during the tournament.
Users traded on almost 60,000 outcomes across both sites during the competition, betting on everything from the sponsor of the Golden Boot winner to whether Cristiano Ronaldo would shed a tear during a Portugal match.
The World Cup, held in the US, Canada and Mexico over June and July, was forecast to be the biggest betting event in history, with a predicted $50 billion in wagers.
Unlike traditional sports betting sites, prediction markets resemble stock exchanges where users trade, via an order book, on whether a real-world event will happen. Prices fluctuate based on what the market believes the probability of that event is. The sites charge fees on each sports trade.
With 48 teams playing 104 games, the World Cup was slated to be the biggest gambling event of all time. Source: Polymarket
The prediction market industry has faced criticism over its vulnerability to insider trading, potential market manipulation and concerns about fueling unregulated gambling. The Wall Street Journal also reported in May that a small number of individuals using algorithmic trading models were taking home an outsized share of winnings.
This would appear to align with Bellingcat’s World Cup analysis, where a small percentage of accounts made most of the winnings. However, the level of detail we were able to obtain did not allow us to see accounts that had utilised algorithmic methods.
Both Polymarket and Kalshi make events and volume data available for programmatic extraction – making it useful for open source analysis. Bellingcat’s data analysis examined all 104 matches as well as the World Cup winner event that was hosted on each platform.
On Polymarket, users traded a total of $10 billion ($5.7 billion on individual games and $4.3 billion on which country would win). The largest game on Polymarket was the Spain vs Argentina final ($212 million), followed by the France vs Spain semi-final ($165 million) and the England vs Argentina semi-final ($142 million).
On Kalshi, users traded a total of more than $4.3 billion ($4.1 billion on the games and $200 million on the winner).
Bellingcat’s analysis also found that 1% of Polymarket trading accounts collected 86% of all winnings during the World Cup, and the bottom 50% of winners shared just 0.1% of profits. The typical winning account on Polymarket made $21, while the typical losing account lost $32 (measured by the median, which is less affected by a handful of exceptionally large wins and losses). More than 12% of traders (14,500) who bet on two or more games lost every bet. The Polymarket account that won the most across all games made a profit of more than $13 million, while the biggest loser lost $11.6 million.
We were unable to run the same win-loss analysis for Kalshi because trading account overviews are not publicly available.
The top teams, by trading volume, across both sites were Argentina ($1.068 billion), Spain ($876 million) and France ($836 million). The top players were Argentina’s Lionel Messi ($40 million), France’s Kylian Mbappé ($36 million) and Norway’s Erling Haaland ($16 million).
How We Calculated the Volume
Polymarket displays the actual traded volume on its site, the total US dollar amount of shares bought and sold since the market started.
Kalshi does not display the traded volume. Instead, it shows the notional volume, which counts every contract traded at the maximum payout value of $1. This means that a token bought for $0.20 will be presented as $1 extra in a user’s displayed volume. This makes the total monetary volume appear higher on Kalshi’s website. To achieve a fair comparison between both platforms, we implemented a heuristic to reconstruct Kalshi’s markets’ volume. We used the daily average price for each market over their duration and multiplied it by the number of contracts traded on that day, the sum of which gives us the values used in this piece. We applied this formula for the more than 21,000 World Cup markets.
Bellingcat is a non-profit and the ability to carry out our work is dependent on the kind support of individual donors. If you would like to support our work, you can do so here. You can also subscribe to our Patreon channel here. Subscribe to our Newsletter and follow us on Bluesky here, Instagram here, Reddit here and YouTube here.
Between February 2022 and September 2025, Bellingcat staff and volunteers collected, geolocated, and shared more than 2,500 incidents of civilian harm following Russia’s full-scale invasion of Ukraine.
As part of this effort, Bellingcat tested a new machine learning model intended to rank Telegram social media posts on their likelihood of containing incidents of civilian harm.
This novel methodology dramatically reduced the search and selection time required, freeing researchers to focus
Between February 2022 and September 2025, Bellingcat staff and volunteers collected, geolocated, and shared more than 2,500 incidents of civilian harm following Russia’s full-scale invasion of Ukraine.
As part of this effort, Bellingcat tested a new machine learning model intended to rank Telegram social media posts on their likelihood of containing incidents of civilian harm.
This novel methodology dramatically reduced the search and selection time required, freeing researchers to focus on verifying incidents of civilian harm – not just searching for them.
This piece documents our methodology, ethical considerations and lessons learned in the hope that others researching similar topics can benefit from our work.
Open source research into civilian harm is still a relatively new field and it presents many challenges – one of the biggest is organising and sorting through the huge volume of user generated content being produced to find what is relevant.
Machine learning, a form of artificial intelligence that uses algorithms to identify patterns from large amounts of data and make predictions, can make this task more efficient.
With ongoing conflicts involving large amounts of civilian harm occurring in Sudan, and much of the Middle East, this guide aims to offer those covering these conflicts an example of how machine learning can be used to help find and sort incidents. You can also access the Code Notebook for our model here.
We defined “civilian harm” not just as civilian deaths or injuries resulting from armed conflict, but also the broader and delayed effects on civilians from mental trauma, loss of livelihood, displacement, destruction of infrastructure and more. This definition was informed by the Protection of Civilians bookon civilian harm.
Initial Telegram Dataset
Each Telegram post containing civilian harm which had already been manually verified by researchers was used to build an initial dataset of confirmed cases of civilian harm, which data scientists call positive instances. We collected a total of 5,848 unique URLs for these Telegram posts. For our manual collection we reviewed posts on relevant Telegram channels, working through oldest to newest posts each day. Assuming that a given post made it to our geolocated incidents list, it meant the researcher who flagged it also looked at the posts that appeared before and after it on Telegram and did not flag those ones, so we selected the 10 posts surrounding the verified civilian harm post as our additional dataset of posts that did not contain civilian harm. After excluding any deleted or duplicate posts, we ended up with 48,545 non-civilian harm posts, our negative instances.
Support Bellingcat
Your donations directly contribute to our ability to publish groundbreaking investigations and uncover wrongdoing around the world.
The choice to overrepresent negative instances aims at better reflecting the real world and increasing data available for model training.
We enriched each URL with metadata from the Telegram API, such as the time of publication, reactions or textual content. As some of these posts had been deleted, we completed the missing data points with previously preserved versions from our Auto Archiver database, only available for the positive instances.
Feature Engineering
Training a machine learning model requires numerical data, as these models compute a prediction score based on mathematical operations.
We built these by converting raw data from our initial dataset, such as keywords signalling potential civilian harm, into numerical scores (or “features”) that the model could interpret, with the aim of increasing the model’s ability to identify patterns. This process, known as feature engineering, can significantly improve model results because it allows data scientists to suggest explicit context knowledge.
A full list of features we used to train the model can be found in the code notebook accompanying this piece. Many features were directly inspired by researchers’ input from their experiences manually screening cases of civilian harm by sorting through a set number of Telegram channels and inspecting each post individually.
Several of the features used were directly built from the metadata contained in each Telegram post including media_type, day_of_week; or binary ones: forwarded, edited andreply_to.
Other features included engagement information: views, forwards, total_reactions, and even individual features for most used emojis including the reaction_crying_face to count emoji.
Converting Text to Numbers
To embed the experience from the manual collection process, researchers put together a list of keywords both in Ukrainian and Russian that, to them, signalled posts likely to show civilian harm. For instance, “Шахед” and “КАБ” translated to “Shahed” and “Guided aerial bomb” respectively. We created a numerical feature to count their frequency.
In addition, we included several generic English-language keywords which meaningfully signalled potential civilian harm, such as “injured”, “school affected” and “hospital affected” that were only used for generating semantic similarity scores.
A semantic similarity score is a calculation used to determine the proximity in meaning between different words and phrases. To get the semantic similarity between the post text and each of our keywords, we represented each in a list of numbers via a Sentence Transformer model, which converts words into numerical representations called vectors that a computer can understand.
We then calculated the level of similarity between each vector using cosine similarity, one of the most popular methods for measuring similarity between two pieces of text.
Due to how embeddings work, this calculation results in a figure on a scale from -1 (no semantic proximity) to 1 (same meaning). For example, the words “hurt” and “injured” would have a high similarity score, while “residential” and “injured” would have a negative score as the words are not semantically similar.
Finally, to enable the model to identify the relevance of each post to civilian harm in Ukraine, we used a multilingual text transformer from the BERT family of language models to represent the entire post’s text as a vector of 768 numerical values. This model can efficiently represent text from many languages in a way that captures meaning: the same sentence in different languages will generate similar embeddings, and trained machine learning models can detect patterns in the embeddings.
It is important to note that for this initial prototype of a civilian harm detection model, we did not include any features derived from media content such as photos and videos, although that would be a logical next step in attempting to improve model performance.
Selecting, Training and Evaluating Models
With 54,393 rows of 893 numerical features each, we selected four machine learning algorithms to train our predictive models.
We chose Logistic Regression as a baseline algorithm due to its simplicity. We also selected three other “best in class” models, Random Forest, XGBoost, and LightGBM. These choices centred on the interpretability of the models and their ability to work on tabular data of this size. For example, we avoided neural networks due to a lack of interpretability and because those models work best with a larger dataset.
To genuinely assess the performance of the trained models, we split our dataset into three parts:
A training set – the data the models were trained on (60 percent of the full dataset’s rows)
A validation set – used for an intermediary evaluation when tuning model parameters (20 percent of all rows)
A test set – hidden for the final performance assessment, so the models were evaluated on unseen data (remaining 20 percent of rows)
We used a stratified split to divide the dataset instead of a random split. This method ensured the proportion of positive instances (i.e. confirmed cases of civilian harm) remained consistent across all three sets at about 11 percent.
To measure the performance of machine learning models, we ran them through the test set and measured the number of correct and incorrect predictions. Models output a likelihood between 0 and 1 that each Telegram post contains civilian harm, and we tried to find a cut-off threshold that leads to a good balance between flagging almost every post (0.1) or flagging very few (0.9).
There are two main types of evaluation metrics to gauge a model’s prediction power. Recall asserts what fraction of positive instances (i.e. known civilian harm posts) were correctly flagged as such. Precision measures the fraction of posts flagged as civilian harm that are indeed civilian harm posts.
During the training phase, we tuned the models to maximise average precision (PR-AUC), a metric that summarises precision across all recall levels. While this method also accounts for precision, it prioritises recall, which is preferable for this use case as it steers model selection to reduce the number of civilian harm posts that are skipped.
The following table sorts models from best to worst PR-AUC against a baseline of a coin-flip predictor. ROC-AUC and F1 are two other evaluation metrics included as sanity checks. Simply put, ROC-AUC measures the probability of ranking two instances, one negative and one positive, correctly; F1 balances precision and recall equally and its best cut-off threshold value.
Model test scores comparison, XGBoost stands out in every relevant metric evaluated.
From these results, we selected XGBoost as our final model as it had the best scores when compared across all metrics.
Interpreting the Model
Because these models are interpretable, we can understand which features are the most useful when predicting whether a post includes civilian harm. The above table shows the top 10 features that most strongly signal the XGBoost model to make a decision:
semantic_keywords_similarity: the semantic proximity between the post text and manually selected keywords “casualties”, “damage” and “civilian harm”
bert: the model was able to discern meaning from the text with the same strength as some of the other features in this list – there are three cases of this in the top 10
reaction_crying_face: reactions with crying face emojis on the post
group_of_messages: whether a post contains multiple media files
keywords_in_text: the number of custom Ukrainian or Russian keywords in the post
These results generally tally with what you might expect when selecting Telegram posts for instances of civilian harm, including that posts that generate a lot of emotional engagement and posts using keywords about civilian harm were among those most likely to contain content related to this topic. Not all models had the same top features as XGBoost. In fact, for the Random Forest model the most important feature was the number of crying face emojis present in a post, a soft pattern highlighted by researchers when this methodology was first imagined.
LLM Results and Comparison
Retroactively, we decided to run a sample of the same test dataset through different large language models (LLMs) to gauge their ability to make these same predictions.
We aimed to include an LLM-generated score as an extra feature for our trained models, which would be captured as relevant if it correlated with the correct predictions.
To start, we selected two local models, the 1B and 4B variants of Gemma 3 from Google DeepMind, and two cloud-hosted models, Gemini 2.5 flash and Gemini 3.5 flash. With this selection, we hoped to compare results across a wide range of models’ expected performance.
We generated a 400-row stratified sample (preserving the same proportion of real civilian harm instances) from the test dataset used for the custom models. For each of the four LLM models, we ran two tests: one where only the Telegram post message was sent, and another including both the message and the engineered features (excluding the text embeddings, as the model had direct access to the text). In the prompt for each model, we asked for a score between 0 and 1. We then evaluated the results as we did for the custom models.
The above table shows that LLMs can indeed extract value from the engineered features. All four LLMs surpassed the baseline Logistic Regression model in our tests, yet none of them performed better than the other custom-trained models, and XGBoost remained the one with the highest PR-AUC.
Still, Gemini 2.5 Flash performed better than its newer version 3.5 and even achieved a slightly higher best F1 score than any other model. While this is a good result, for the flagging of civilian harm posts, the PR-AUC remains the crucial metric, as it captures the model’s ability to identify infrequent instances of civilian harm while minimising false positives.
Ethical Considerations
Introducing an instrument of automated decision-making into a process of detecting civilian harm brings inherent ethical questions. These include automation bias, or how humans tend to blindly place faith in machine-generated recommendations; algorithmic bias, or how the results of these models echo the same patterns present in the training data, including under- or over-representation of types of civilian harm.
The decision to test an automated methodology for this particular project came from the fact that there were limited resources for both steps in the process – the detection of potential civilian harm and its actual verification. Historically, we built an enormous backlog of unverified incidents because a lot of time had to be spent on monitoring the most recent events so that potential evidence would be captured and preserved as soon as possible.
The automation of this process also reduced the exposure of researchers to a significant amount of unpleasant and distressing visual and text content, reducing the burden of exposure to traumatic content.
For this project, we tried to ameliorate the ethical challenges with a number of strategies including randomly flagging posts not captured by any model, monitoring which features models relied on to make decisions, and by doing historical comparisons of patterns in data.
Additionally, as stated above, for this initial prototype of a civilian harm detection model we did not include any features derived from the media content itself. In the future, it would be a logical next step in attempting to improve the model performance, to include the media from the posts – but using AI to review actual media comes with additional ethical challenges such as model bias.
Because of the opaque ownership of many LLM companies and their generative nature, the use of LLMs for an extra feature presented additional ethical challenges including privacy and safety concerns considering the sensitive nature of the data. Our model did not rely on LLMs, though we retroactively ran a sample through it.
How the Model Fits into the Bigger Picture
After selecting this model, we created a user interface where researchers could view a list of Telegram posts sorted from most to least likely to contain indications of civilian harm. The user interface was designed for quick triage and integration, where a positive confirmation from researchers would instantly send the post to the Auto Archiver (Bellingcat’s tool for preserving digital content) and then transfer it to ATLOS (our internal collaborative verification platform). Bellingcat staff and volunteers could then manually verify incidents. Researcher input was constantly stored so that this data could be used to improve the model in the future.
Preliminary feedback indicated that the AI model was useful. Not only were we able to reduce time and harm from scouring through dozens of war reporting Telegram channels, researchers also reported that the stream of new posts being added to the verification backlog were capturing real and diverse cases of civilian harm.
We recognise this model has much room for improvement and is a work in progress. Even though it can illicit diverse civilian harm posts, further tests and improvements (such as improved feature engineering and continuous evaluation) are needed before it can confidently be deployed.
Despite the focus on civilian harm and Telegram (highly popular in Ukraine and Russia), this pipeline is generic and can be adapted to other conflict monitoring tasks. How easily this can be done does depend on how open the social media platform is and whether it is possible to scrape posts from it. Apart from that, it is easy to incorporate new features and data, and cheap to automatically retrain, test and deploy models as the system receives more human input.
Looking forward, sorting through overwhelming amounts of data in a conflict will continue to be challenging. Hopefully, this methodology can help newsrooms, conflict monitoring organisations, and others find the balance between ethical considerations and resources in order to carry out open source investigations on civilian harm and human rights violations.
Editor’s note: This article was updated on July 3, 2026, to include a line outlining that the model described is a work in progress.
Bellingcat is a non-profit and the ability to carry out our work is dependent on the kind support of individual donors. If you would like to support our work, you can do so here. You can also subscribe to our Patreon channel here. Subscribe to our Newsletter and follow us on Bluesky here, Instagram here, Reddit here and YouTube here.
Support Bellingcat
Your donations directly contribute to our ability to publish groundbreaking investigations and uncover wrongdoing around the world.
1intro
2WHO ARE THE RUSSIAN
COSSACKS?
3EVERYTHING THROUGH YOUTH
4FROM WAR GAMES TO REAL WEAPONS
5VOLUNTARY RECRUITMENT
6BARS-15
7IN THE END
Your browser does not support the video tag.
From School to Battlefield to Grave
How Russian Cossacks drive young people to war
This video was
post
This is
Олег Монин
who took Berkut’s oath
four months earlier. Through this veiled Cossack
Youth Organisation, he trained in combat tactics with returned fighters and transitioned from
pretend to real weapons.
Within a year, Oleg abandoned his studies and enlisted in
БАРС-15,
a Cossack Volunteer Battalion fighting in Ukraine.
By Feb. 10, 2025 Oleg was dead. He died aged 19, less than four months after deployment in Ukraine.
Cossack societies, organisations, and even military units provide an identity that is indigenous to
Russia, Visiting Assistant Professor at Miami University, Dr Marcello Fantoni told Bellingcat.
This
identity is “rooted in ‘traditional’ values, martial prowess, military readiness, orthodox
religiosity and a culture not influenced by the ‘corrupting’ West,” Fantoni added via email. This is
why “education is central to the overall enterprise”.
Oleg’s story demonstrates how the Cossacks drive young people from a school club to a war zone and
enable a state-sponsored alternative mobilisation force.
WHO ARE THE RUSSIAN COSSACKS?
The Cossacks played an important role in the formation of the Russian Empire. They lived in
communities called hosts on the edges of the empire. They operate under
a military hierarchy ruled by a chief,
the Ataman. Due to their loyalty to the Tsar, the Cossacks were repressed by the
Bolsheviks after 1917.
Credit: Journal “Chronicle of War”, 1915; Nicholas II among officers
When the Soviet Union collapsed in 1991, the Cossacks’ descendants
called for a “rebirth”.
In 2005, a bill submitted by President Vladimir Putin allowed registered Cossack organisations members
to serve in military units and police forces.
Credit: tamvesti.ru
New hosts were created in traditionally non-Cossack lands with a
variety of institutions
to direct them. In 2018, the government
united them in the “All-Russian Cossack
Society”. Putin tries to marginalise the traditional Cossack groups, analyst Paul Goble told
Bellingcat while the ones “he has created for his own purposes” play a “major role in military and
patriotic education”.
Credit: Kremlin
Only 8 of Russia’s 83 recognized
Federal Subjects do not
have a registered Cossack Host.
In 2018, the Black Sea Cossack Host of Crimea
entered
the register.
The peninsula has been under Russian occupation since 2014. The Cossack legacy is also vitally important to Ukrainian
identity.
There are new
hosts in the occupied Ukrainian territories of Kherson, Zaporizhzhia, Donetsk, and
Luhansk.
Russian Cossack organisations have been “very active within the occupied Ukrainian regions,” Dr
Fantoni told Bellingcat. They “recruit local residents and then deploy them for cultural and
military purposes,” allowing Russia “to contest and even co-opt a central tenet of Ukrainian
national identity – Cossackdom,” he said.
The national “All-Russian Cossack Society” VSKO was
created
in 2018,
and in 2019, the
State
Duma
gave Russian President Vladimir Putin exclusive authority to appoint its national Ataman.
At the top of the VSKO is Ataman Vitaly
Kuznetsov, a Cossack General.
Kuznetsov was
appointed
in November 2023, succeeding the first-ever national Ataman – Nikolai Doluda, then 70 years old
and a
sanctioned individual.
Kuznetsov has also become a leading Cossack interacting with the Russian state.
Including with Dmitry Mironov,
assistant to President Putin and Chair of the Council for Cossack Affairs.
As well as Leonid Pasechnik, head of the Luhansk People’s Republic. Kuznetsov
thanked
Pasechnik in June for helping create three Cossack Cadet Corps in the occupied region.
According to Kuznetsov, the VSKO
priorities are “development of military Cossack societies in all directions: education, culture,
history, and most importantly, youth. Everything through youth.”
EVERYTHING THROUGH YOUTH
Cossack education can be divided into primary, secondary, and tertiary levels, all with the goal
of
promoting a unified system.
There are Cossack schools and regular schools with a Cossack affiliation.
Data from 2022
claim there were just under 2000 such institutions with around 210,000 students, but
recent claims point to over 300,000
students.
Oleg’s story demonstrates how young people outside formal Cossack education can still get pulled
in. It also shows that the Cossacks are but one of several interlaced strategies for
“military-patriotic” education.
Oleg grew up in Saratov.
Credit: Image of youth practicing putting on a gas mask, posted on VKontakte by Lyceum N.3.
He studied in Lyceum N.3, a state-funded educational institution in Saratov. Often, the school
promotes events like the national
Zarnitsa
competition. It includes activities like “putting on gas masks” or “sniper games” for third
graders.
The school’s
military club “Fakel”
acts as an intermediary for these events and other nationwide military education initiatives
such as the 24-hour-long
Avangard training for tenth graders.
In 2024,
Natalia
Saprykina,
the director of Lyceum N.3, was
awarded
a Letter of Gratitude for her “contribution to the patriotic education of the younger
generation” by
a Deputy of the
Regional
Duma.
Oleg graduated from high school in 2023
at the age of 17.
In the same year he enrolled in InPIT, a
higher education
institution of the Saratov State Technical University.
By November Oleg had turned 18 and was
wearing
military fatigues and practising survival skills alongside other candidates of a
“military-patriotic” student association named Berkut, at another local university, the
Saratov State Law Academy (SSLA).
Though Berkut is not explicitly a Cossack organisation, we established several
connections between the head of Berkut, Alexander Andreevich, and Cossack organisations. As
we’ll see, Andreevich was present at multiple military style training camps that Oleg took
part in.
Neither Berkut’s VKontakte nor Telegram channel descriptions mention the Cossacks.
The association’s official objectives are “forming a positive image of military service” and
“popularisation of service in the Russian army and law enforcement agencies”. It is headed by
Alexander Andreevich.
However, some of Berkut’s
videos
include the banner of a
Молодёжная казачья
организация.
A Telegram
post
by Andrey Fetisov, the Saratov District Ataman, refers to Berkut as a “Cossack Youth Movement”.
Even though Berkut (left) shares a name and eagle iconography with a notorious
Ukrainian special police
force (right), part of which defected to Russia during the occupation of Crimea in 2014,
Bellingcat found no link between the two organisations.
FROM WAR GAMES TO REAL WEAPONS
By December 2023, nearing the end of the first semester, Oleg and the other candidates
took the Berkut oath,
making them official members. Oath-taking ceremonies are
“invented
traditions”
among Cossack forces.
Atop the dais stand senior members of Berkut, including the head of the organisation – Alexander
Andreevich.
Andreevich is an active Cossack who has been working under the guidance of District Ataman
Andrey Fetisov since at least April 2023.
More recently, in January 2025, they were both delivering a
lesson
to Cossack children for Yunarmiya,
exemplifying the overlapping network of youth militarisation initiatives.
In August 2024, Andreevich attended the
iVolga
Cossack Youth Festival, where he met
Kuznetsov.
The only two people featured speaking in an
official video.
Andreevich also led Oleg to two military-inspired events in April 2024.
Five days later they went to a training that
included trench tactics and simulated helicopter jumps.
Since
2023,
Oleg often wore a distinctive yellow and red
“Скорпион”
call sign patch on his chest when wearing military fatigues, which distinguishes him from other
youth at the events. That and other distinctive features identify him even with a mask or
goggles.
Bellingcat was able to geolocate this place to be a
Rosgvardia
training ground
on the outskirts of Saratov.
Notably, the trenches are not visible on Google Earth but are on Yandex Maps, which has more recent imagery for
the region.
As is
Oleg Mysov, another returned fighter who also
engages in “patriotic education of youth”
events.
Both have
attended Cossack events.
Even though in this photo they are holding the Volga Cossack Host flag, Bellingcat could not
clearly identify them as Cossacks.
Bellingcat geolocated it to a military training ground in Samara, the same location where
other
Cossack recruits
trained before deploying to BARS-15. Fetisov himself
shared photos
of this training ground
two weeks after stepping down
as Ataman to join BARS-15. Andreevich left and Oleg right in
this
photo.
They used real weapons this time. A
video montage
shows participants firing live rounds.
This is
a
photo
that includes Oleg, Fetisov, and Andreevich. The first media we found for this event
is from early September
which is consistent with the
sun position
in this photo and the grass patches seen in satellite imagery from early September 2024.
Bellingcat contacted Kuznetsov, Fetisov and Andreevich to ask about their roles in the Cossack
community, but they haven’t responded.
This is the last time Bellingcat was able to trace Oleg’s whereabouts with open sources before
he joined BARS-15.
VOLUNTARY RECRUITMENT
Many countries have a volunteer reserve system for getting more soldiers in times of war. In
Russia, the system is known as
BARS,
created in 2015
and
intensified
in
2021.
All BARS fighters sign a contract with the Ministry of Defense and get paid.
Mapping the geolocated positions of these units in the
UAControlMaps Project dataset
reveal widespread areas of operations. BARS Battalions are often reorganised.
Estimates
put the total number so far at
over 30 BARS Battalions and 10 of them
have overt Cossack affiliation.
Cossacks also
operate
as detachments in other military structures.
By
their own reckoning,
in February there were more than 18,500 Cossacks on the front lines in Ukraine. In May the
first-ever national Ataman, Nikolai Doluda, gave a higher figure of 46,000
Cossacks.
As of 2024, British Professor Rod Thornton estimated
that BARS constitute some 10-30,000 troops in Ukraine, 15% of the total invasion force.
The
Mediazona project
tracks individual Russian losses in Ukraine and publishes bi-weekly reports. As of Nov. 21, 2025,
they identified 149,241 publicly named casualties,
Oleg
among them.
The project also tracks volunteer casualties.
Deaths of volunteer fighters constituted 12.8% of losses in 2022 and 21.9% 2023. In 2024 they more
than doubled to 45.7%. As of Nov. 21, verified
deaths of volunteer fighters for 2025 were at 42.8%.
BARS-15
BARS-15 is a Cossack battalion created on
May 15, 2022,
and named
Ермакafter
a
historical
Ataman.
Originally composed of Cossacks from multiple hosts, mainly Volga and Oremburg, it
now
draws its members from the Volga Host only.
Credit: All-Russian Cossack Society
The panel reads Black Hussars. Oleg is on his knee in front of Andreevich, wearing his distinctive
“Scorpion” patch.
Credit: VKontakte @svpo_berkut
The number of active Cossack fighters in BARS-15 is reportedly 400, a number echoed
by a former Commander, with other sources saying over 900 volunteers
have passed through as of September 2024. They reportedly
took part in the
invasion
of Avdiivka among other combat
activities in Ukrainian cities both in Donetsk and Luhansk.
Credit: VKontakte @vvko_russia
One of its former members is Andrey Fetisov, who temporarily stepped down as Saratov District Ataman and joined
BARS-15 between approximately November
2023 and June 2024.
Credit: Telegram @izvestia64
In April 2024, Fetisov received a Medal for Bravery from
Vitaly Kuznetsov, the national Ataman. Within six months, Fetisov would be taking Oleg to the
BARS-15 training camp.
Credit: Telegram @izvestia64
There are many reasons why people are motivated to join Cossack groups, Dr Fantoni told Bellingcat,
adding that these motivated individuals “are the driving force” behind militarisation. “Some do it
out of patriotic motivations, others for political, economic or individual status gain, some even
because this can protect oneself from future mobilisation to an actual fighting unit,” he said.
This image first appeared on Oleg’s obituary
posted
by Fetisov. The vehicle, road, and equipment are consistent with those used by other fighters with
the
Black Hussars
around February 2025.
According to
recruitmentposts
BARS-15 training takes three weeks. A
recent
study
found that to be the norm in Russia’s military while also labelling training as “low-quality and
ineffective”.
Bellingcat reached out to Oleg’s parents.
His mother said she couldn’t speak about Oleg’s death,
it still hurts too much.
Additional research by Timothy B, Afton Briones, Sarah Grossman, Alexandra Malikova, Mitchell Polman, Olivia
Gresham, Bonny Albo, Adam Arthur, Robert Chapman of the Bellingcat Volunteer
Community.
Youri van der Weide and Aiganysh Aidarbekova contributed to this report.
Bellingcat is a non-profit and the ability to carry out our work is dependent on the kind support of
individual donors. If you would like to support our work, you can do so here.
You can also subscribe to our Patreon channel here.
Subscribe to our Newsletter
and follow us on Bluesky here and Mastodon here. With the
unpredictability of social media algorithms making it harder for news outlets to reach audiences
consistently, we have also started a WhatsApp channel that you can join to stay updated on our
stories.
Satellite images are courtesy of Yandex, Maxar, Airbus, MapBox and Google Earth.
Co-funded by the European Union. Views and opinions expressed are those of the author(s) only and do not
necessarily reflect those of the European Union or the European Health and Digital Executive Agency
(HADEA). Neither the European Union nor the granting authority can be held responsible for them.