Methodology
How BotShark decides a score
BotShark looks at public posting behaviour and asks a practical question: does this account look automated, coordinated, or farmed? The answer is a 0–100 likelihood score, a plain-language tier, and a report that shows why — not a black-box “bot / not bot” label.
What we measure (and what we don’t)
We measure behavioural likelihood — posting rhythm, engagement shape, content reuse, profile cues, and network patterns that commonly show up on automated or coordinated accounts.
We do not score ideology, “truthfulness,” or guilt. A high score means several independent signals lined up with automation or coordination heuristics. You still decide what that means in context.
Every point on the score maps to a published rule and, where possible, the posts or metrics behind it. If we can’t show the work, we don’t pretend the evidence is stronger than it is.
The pipeline
-
01
Collect
Public profile metadata and recent posts/replies from supported platforms (X, Bluesky).
-
02
Extract
Behavioural features: cadence, engagement asymmetry, reuse, thread shape, network samples.
-
03
Score
Explainable rules add points when their conditions match. Overlapping rules don’t double-count.
-
04
Explain
The report lists what fired, in plain language, with evidence — plus confidence about data coverage.
What goes into the score
BotShark is not a machine-learning classifier that “feels” bot-like. It’s a stack of inspectable rules. Each rule asks a specific question about public behaviour. If it matches, it adds points (some rules only annotate evidence and add zero points). The score is the sum, capped at 100.
How often they post
Volume, bursts, regularity, and whether posting looks like a schedule instead of a person.
- Bursts of many posts within minutes
- Shows up to post at almost all hours
- Gaps between posts are oddly regular
How they interact with others
Reply farms, one-way engagement, template replies, and broadcast-only behaviour.
- Replies out constantly, gets almost nothing back
- Groups of near-identical replies
- Mostly broadcasting, not conversing
Profile & bio
Default avatars, digit-stuffed handles, promo bios, and bought/repurposed account cues.
- Handle ends in a long run of digits
- Sat quiet for ages, then suddenly active
- Location doesn’t match posting hours
What they publish
Repeated text, link spam, salesy language, and unnatural stylistic sameness.
- Same text keeps showing up
- Almost every post includes a link
- Writing style barely changes
Followers & following
Lopsided graphs, follow-for-follow growth, and cohort patterns in who surrounds the account.
- Follows many, followed by few
- Follower and following counts badly lopsided
When they’re usually online
Timezone and waking-hours checks — especially when a claimed location doesn’t match activity.
- Says they’re American, posts on another schedule
How their replies read
LLM-polished or oddly generic reply texture when enough reply text is available.
- Lots of generic one-liner replies
Promo, scam, and coordination cues
Grift funnels, recycled sockpuppets, and same-words-as-other-scanned-accounts signals when corpus context exists.
- Bio reads like a promo or scam
- Same words as other accounts you’ve scanned
Example headlines above are the same plain-language lines used in live reports. Platform packs (X, Bluesky, …) share the core idea and adapt collection to each surface.
How points become a number
- Additive rules. Matching rules contribute points. Stronger or more extreme patterns can contribute more.
- No double-counting the same story. When two rules describe the same underlying behaviour, we keep the stronger one so the score doesn’t inflate.
- Cap at 100. The score never exceeds 100 even if many rules fire.
- Confidence is separate. Confidence reflects how complete the sample was (empty timelines, missing network samples, shallow vs deeper thread analysis) — it is not folded into the bot score.
Likelihood tiers
The number maps to a tier so readers don’t have to memorise cutoffs. Tiers describe likelihood of automation/coordination patterns, not a courtroom verdict.
| Score | Tier | How to read it |
|---|---|---|
| 0–34 | Low likelihood | In what we sampled, this mostly behaves like a normal user. That doesn’t mean they’re trustworthy — just that we didn’t see strong bot patterns. |
| 35–59 | Moderate likelihood | A mix of habits that raise an eyebrow and some that look ordinary. Read the details rather than going on the score alone. |
| 60–79 | High likelihood | Several parts of the account line up with how bots and amplifier farms usually behave. No single number proves it, but the overall shape is suspicious. |
| 80–100 | Very high likelihood | Posting habits, engagement shape, and content all tell the same story — this doesn’t read like someone using the app for themselves. |
What’s in a report
A finished report includes the score and tier, the strongest signals in plain language, category breakdowns, charts (posting rhythm, engagement mix, and related views), and caveats when the sample was thin. Optional deeper modes expand thread and follower sampling when available.
Treat the report as one investigative input alongside primary sources and human judgement. See neutrality before citing findings publicly.
Limits (read these)
- Empty timelines hide bots. Some automated accounts have little or no visible posts in the window we can collect. We can’t invent posting patterns that aren’t there — those cases need network/history cues or a deeper rescan, not a fake high score.
- Public data only. Private accounts, deleted posts, and platform rate limits bound what we see.
- Humans can look weird; bots can look polished. A low score is not a character reference. A high score is not proof of intent.
- We revise rules when the data disagrees. Thresholds and weights are tested against labelled known-bot and known-human accounts, then updated when a rule overfires or underfires.
How we check the scoring
We keep a labelled set of accounts we believe are bots and accounts we believe are humans, run BotShark on them, and ask two everyday questions at each score cutoff:
- When we call something a bot, how often are we right?
- (Researchers call this precision.) At score ≥ 60 on accounts with posts we could score , that was 100% (roughly 85–100% of the time).
- Of the known bots in that set, how many do we catch?
- (Researchers call this recall.) Same cutoff : 67% (roughly 50–80%).
Current labelled set: 166 accounts (123 bots, 43 humans). 90 had empty timelines when collected · eval refreshed 14 July 2026. The ranges in the tables below are 95% confidence intervals — a statistical “we’re not pretending one sample is exact.”
Accounts with posts we could score
The set we recommend citing. Skips labelled bots that had an empty timeline when collected — those cannot show posting patterns, so including them makes “catch rate” look worse than it is.
| If score ≥ | Right when we say bot | Known bots caught | Balance (F1) |
|---|---|---|---|
| 40 | 73% (56–85%) | 73% (56–85%) | 0.73 |
| 50 | 88% (71–96%) | 70% (53–83%) | 0.78 |
| 60 | 100% (85–100%) | 67% (50–80%) | 0.80 |
| 70 | 100% (84–100%) | 61% (44–75%) | 0.75 |
| 80 | 100% (83–100%) | 58% (41–73%) | 0.73 |
Accounts that never looked empty
A stricter slice: labelled accounts whose timelines had content throughout collection.
| If score ≥ | Right when we say bot | Known bots caught | Balance (F1) |
|---|---|---|---|
| 40 | 73% (56–85%) | 73% (56–85%) | 0.73 |
| 50 | 88% (71–96%) | 70% (53–83%) | 0.78 |
| 60 | 100% (85–100%) | 67% (50–80%) | 0.80 |
| 70 | 100% (84–100%) | 61% (44–75%) | 0.75 |
| 80 | 100% (83–100%) | 58% (41–73%) | 0.73 |
Every labelled account
Includes empty-timeline bots. Useful for honesty about coverage gaps — not the number to quote as “how good is BotShark at catching bots.”
| If score ≥ | Right when we say bot | Known bots caught | Balance (F1) |
|---|---|---|---|
| 40 | 74% (57–85%) | 20% (14–28%) | 0.32 |
| 50 | 89% (72–96%) | 20% (14–27%) | 0.32 |
| 60 | 100% (85–100%) | 18% (12–26%) | 0.30 |
| 70 | 100% (84–100%) | 16% (11–24%) | 0.28 |
| 80 | 100% (83–100%) | 15% (10–23%) | 0.27 |
These tables package measurements from our labelled eval runs. They do not invent new claims. For a walkthrough of the product surface, start with how it works .