M37

What “100% detection” counts

This is the material behind the video. MITRE’s ATT&CK Evaluations are the one head-to-head test of endpoint detection and response products that publishes its whole per-product record afterwards, and in the 2023 Turla round twenty-nine of them were run against the same scripted attack, broken into 143 scored substeps. Five detected every one. That claim is true, and the record grants it.

The same record also says what kind of detection each one was. Every result carries a label: the product named the technique, or named the goal without the method, or raised an unclassified alert, or recorded the raw event and said nothing about it. On that second axis the same field, on the same attack, runs from 99.3% down to 23.8%.

The two products the second axis separates furthest are neighbours on the first. Elastic detected 80.4% of the substeps and named the technique on 31.5% of them. Deep Instinct detected 82.5% and named the technique on 79.0%. That is 2.1 points apart on the number that gets quoted and 47.5 points apart on the one that does not. Neither figure is wrong. They are answers to different questions, and only one of the questions is usually asked.

How this was counted

The method comes before the results because on this page it is the whole difference between a true statement and a misleading one. Both come out of the same published file.

The scenario this page does not count

A round’s scenarios are not all detection tests. In this round the first two are the detection evaluation. The third re-runs the same attack as a protection test, which asks whether the product blocks the activity rather than whether it reports it. Detections are not scored there at all, so every cell in it reads “None” for every product.

Counting it produces a headline that follows correctly from the raw file and is wrong about detection: 131 of 274 substeps missed by all 29 products, 47.8 per cent. That is the number this page nearly carried, and it is the most useful thing on it. Nobody was being asked to detect anything in those 131 substeps.

The exclusion is made mechanically rather than by writing the scenario’s name into the code: a scenario in which not one product recorded a single detection is not a detection test, whatever it is called, so a round that renamed or reordered its scenarios would still be excluded correctly. It is corroborated a second way. The excluded scenario’s substep criteria are a 100% match for the scored scenarios’, because it is the same attack run again. Both of those are asserted when this page loads its data, and if either stops holding the page fails to build rather than publishing the wrong number.

What this page does not see

It reads what MITRE published and nothing else. It does not see whether a detection would fire outside a lab, on a network with its own noise and its own exceptions. It does not see whether anyone at the console noticed the alert or did anything about it. It does not see how a product was configured or tuned before the run, beyond the record’s own note on the cells where a detection followed a configuration change made during the evaluation. And it says nothing whatever about a product that did not take part: a vendor absent from a round is absent, not worse.

The rounds are not a series

The last table on this page lists all 7 evaluation rounds. They are different exams rather than points on a line. The adversary changes between them, the number of substeps changes, the roster of participants changes, and so does the scoring vocabulary. One product’s technique-level share in one round and another product’s in another are answers to different questions, and the difference between them is not progress.

The 2018 APT3 round is the clearest case. It labelled detections with a different vocabulary altogether — General Behavior and Specific Behavior — so its technique-level figures read as not scored rather than as zero. The products did not fail on that axis; the axis did not exist yet. The same shows up elsewhere in the table wherever a whole rung of the ladder reads zero.

MITRE does not name a winner

The evaluations publish results and do not rank the participants. There is no score, no league table and no pass mark in the evaluation itself; the rankings that circulate afterwards are made by whoever is quoting it, which is usually a vendor. This page does not add one. It will order a table when you click a heading, which is not the same thing as saying which product is better — that depends on what you run, who watches the console, and what they do when it lights up, and none of those are in the record.

Where the numbers came from

Every figure below is a count of cells in the published results, 4,147 of them: one per product per substep, 29 times 143. They were read from the evaluations site by code at commit ffa5cbecadb9, on 2026-09-05. The evaluations have moved host once already, which is why the harvest keeps its own copy of the files it measured.

The ladder

Each of the 4,147 cells carries one of these labels. The first four are the evaluation’s own words for what a detection was; the last two are how this page accounts for the rest.

Technique · 2,492
The product named the technique: not that something happened, but which ATT&CK technique it was.
Tactic · 202
The product named the goal — credential access, lateral movement — without naming the method used to reach it.
General · 323
Something suspicious was flagged, with no classification attached to it.
Telemetry · 425
The raw event was recorded and nothing was said about it. Finding it is the analyst's job.
Anything else · 0
A label outside the four above. There are none in this round, which is itself worth knowing: the vocabulary is not constant between rounds, and it is what makes two rounds hard to compare.
No detection · 705
Nothing was recorded for that substep by that product.

A “detected” percentage adds the first four together and divides by 143. That is a fair thing to count: a telemetry record is a real record, and on a team with people reading the console it may be the one that matters. It is simply not the same measurement as the first line, and the difference between the two is the rest of this page.

The numbers in the video

5 of 29
Products that detected every one of the 143 scored substeps
n = 29
What it does not mean. It does not mean the five are interchangeable. All of them detected every substep, and their technique-level shares still run from 93.0% to 99.3%, which the record states separately for each of them.
23.8% to 99.3%
The technique-level share across the same field, same attack, same substeps
n = 4,147
What it does not mean. It is not a measure of protection. It counts what the record says each detection was called, not whether the activity was stopped, nor whether anybody acted on the alert.
60 of 143
Substeps every one of the 29 products detected
n = 143
What it does not mean. It does not mean those 60 were easy. It means they separate nobody from anybody, and since no substep at all was missed by every product, the 83 remaining substeps carry the whole difference between these products.
20 of 29
Products with identical scores when the round is read at its 19 steps
n = 19
What it does not mean. It does not mean those 20 products behave alike. It means a 19-step summary of this round cannot tell them apart, and a summary is what most people are shown.

The 29 products, on both axes at once

One row per product. Detected is the share of the 143 substeps on which the product produced any detection at all, at any rung of the ladder; named the technique is the share on which the record labels that detection Technique. The four rung columns between them are counts of substeps, and they add up to the detected count. Opening a row shows all of them, the gap between the two percentages, and the two modifiers the record carries alongside a detection.

Showing 25 of 29 products. Sorted by the share named at technique level, largest first.

The second axis on its own, highest first. This is the same field the headline ranks flat, and it is the ordering the headline cannot produce.

Row detail
Palo Alto Networksdetected in 19 of 19 steps100.0%143 of 14399.3%142 of 143142100019 of 19
Cybereasondetected in 19 of 19 steps100.0%143 of 14396.5%138 of 143138221019 of 19
CrowdStrikedetected in 19 of 19 steps100.0%143 of 14395.8%137 of 143137600019 of 19
Fortinetdetected in 19 of 19 steps97.9%140 of 14393.7%134 of 143134024319 of 19
Cynetdetected in 19 of 19 steps100.0%143 of 14393.0%133 of 143133820019 of 19
Microsoftdetected in 19 of 19 steps100.0%143 of 14393.0%133 of 143133433019 of 19
Sophosdetected in 19 of 19 steps98.6%141 of 14379.7%114 of 1431141791219 of 19
Deep Instinctdetected in 18 of 19 steps82.5%118 of 14379.0%113 of 1431130322518 of 19
HarfangLabdetected in 19 of 19 steps87.4%125 of 14376.9%110 of 14311012121819 of 19
SentinelOnedetected in 18 of 19 steps88.1%126 of 14376.9%110 of 1431103761718 of 19
TrendAIdetected in 19 of 19 steps88.1%126 of 14376.2%109 of 1431094581719 of 19
Bitdefenderdetected in 19 of 19 steps91.6%131 of 14375.5%108 of 14310814541219 of 19
Uptycsdetected in 19 of 19 steps88.1%126 of 14373.4%105 of 1431058761719 of 19
ThreatDowndetected in 19 of 19 steps82.5%118 of 14367.8%97 of 1439738102519 of 19
Trellixdetected in 19 of 19 steps81.1%116 of 14356.6%81 of 14381716122719 of 19
ESETdetected in 19 of 19 steps77.6%111 of 14347.6%68 of 14368413263219 of 19
WatchGuarddetected in 18 of 19 steps72.7%104 of 14347.6%68 of 14368222123918 of 19
IBM Securitydetected in 19 of 19 steps72.0%103 of 14346.9%67 of 14367617134019 of 19
BlackBerrydetected in 19 of 19 steps81.8%117 of 14342.7%61 of 14361326272619 of 19
Secureworksdetected in 19 of 19 steps78.3%112 of 14341.3%59 of 14359114383119 of 19
VMware Carbon Blackdetected in 19 of 19 steps72.0%103 of 14340.6%58 of 14358224194019 of 19
Broadcom Symantecdetected in 18 of 19 steps75.5%108 of 14339.9%57 of 143571026153518 of 19
WithSecuredetected in 18 of 19 steps67.1%96 of 14337.1%53 of 14353168194718 of 19
Qualysdetected in 18 of 19 steps78.3%112 of 14333.6%48 of 14348422383118 of 19
Elasticdetected in 19 of 19 steps80.4%115 of 14331.5%45 of 14345530352819 of 19
25 of 29 shown

The 143 substeps, and who caught them

One row per scored substep of the attack, in the order the record publishes them. What the attacker did is the evaluation’s own criteria text, not a summary of it. Detected by counts products, out of 29; named the technique counts the subset of those that identified it rather than only recording it. Every technique id opens its ATT&CK page.

Showing 25 of 143 substeps. In the order the record published them.

Row detail
1.A.1step 1 · Initial Access1Initial AccessT1566 Phishing29of 29165Gunter clicks link in email from noreply@sktlocal.it and downloads NTFVersion.exe
1.A.2step 1 · Execution1ExecutionT1204 User Execution29of 29220Gunter executes NTFVersion.exe
1.A.3step 1 · Defense Evasion1Defense EvasionT1027 Obfuscated Files or Information29of 29172mxs_installer.exe is embedded inside NTFVersion.exe via the Resource Section
1.A.4step 1 · Defense Evasion1Defense EvasionT1112 Modify Registry29of 29251NTFVersion.exe modifies Gunter's Winlogon Registry Key
1.A.5step 1 · Persistence1PersistenceT1547 Boot or Logon Autostart Execution29of 29242NTFVersion.exe modifies Shell key value to include mxs_installer.exe
2.A.1step 2 · Defense Evasion2Defense EvasionT1027 Obfuscated Files or Information19of 29160EPIC Guard DLL is embedded inside the Resource Section of mxs_installer.exe
2.A.10step 2 · Command and Control2Command and ControlT1071 Application Layer Protocol28of 29224msedge.exe connects to shoppingbeach[.]org over HTTP protocol
2.A.11step 2 · Command and Control2Command and ControlT1090 Proxy22of 29142msedge.exe connects to adversary's compromised proxy - shoppingbeach[.]org
2.A.2step 2 · Defense Evasion2Defense EvasionT1055 Process Injection28of 29240mxs_installer.exe injects EPIC GUARD DLL into explorer.exe via CreateRemoteThread
2.A.3step 2 · Discovery2DiscoveryT1057 Process Discovery22of 29162explorer.exe enumerates process list via CreateToolhelp32Snapshot
2.A.4step 2 · Defense Evasion2Defense EvasionT1027 Obfuscated Files or Information13of 29111EPIC Worker DLL is embedded inside the Resource Section of explorer.exe's Guard DLL
2.A.5step 2 · Defense Evasion2Defense EvasionT1055 Process Injection27of 29210explorer.exe injects EPIC Worker DLL into msedge.exe via CreateRemoteThread
2.A.6step 2 · Discovery2DiscoveryT1087 Account Discovery22of 29143msedge.exe enumerates all users on the local machine via NetUserEnum
2.A.7step 2 · Discovery2DiscoveryT1083 File and Directory Discovery23of 29164msedge.exe enumerates Gunter's files via FindFirstFile & FindNextFile​
2.A.8step 2 · Collection2CollectionT1560 Archive Collected Data11of 2982msedge.exe bzip2 compresses discovery output in memory
2.A.9step 2 · Command and Control2Command and ControlT1132 Data Encoding12of 2991msedge.exe base64 encodes discovery output in memory
3.A.1step 3 · Discovery3DiscoveryT1069 Permission Groups Discovery29of 29290cmd.exe executes various net group commands
3.A.2step 3 · Discovery3DiscoveryT1007 System Service Discovery29of 29271cmd.exe executes "tasklist /svc"
3.A.3step 3 · Command and Control3Command and ControlT1573 Encrypted Channel13of 29100msedge.exe uses a temporary AES key to encrypt command output in memory
3.A.4step 3 · Command and Control3Command and ControlT1573 Encrypted Channel14of 29102msedge.exe RSA encrypts the AES temporary key in memory
3.A.5step 3 · Discovery3DiscoveryT1012 Query Registry29of 29290cmd.exe reg queries the ViperVPNSvc service
3.A.6step 3 · Execution3ExecutionT1059 Command and Scripting Interpreter29of 29252cmd.exe executes powershell.exe to verify to which users can access the ViperVPN service
3.A.7step 3 · Privilege Escalation3Privilege EscalationT1574 Hijack Execution Flow29of 29212cmd.exe modifies the ViperVPN service registry key
4.A.1step 4 · Command and Control4Command and ControlT1105 Ingress Tool Transfer29of 29206svchost.exe creates C:\Windows\System32\WinResSvc.exe
4.A.10step 4 · Command and Control4Command and ControlT1573 Encrypted Channel15of 29101msedge.exe receives CAST-128 encrypted tasking
25 of 143 shown

The rounds, and why they are not a series

Every round the evaluations have run, oldest first, with the one this page is about marked. Read down a column and you are comparing adversaries, rosters and scoring vocabularies as much as products. The reasons are above; the table is here because leaving it out would let one round read as the state of the art rather than as one exam of it.

YearRoundProductsSubstepsTechnique-level shareDetected everything
2018APT312136not scored on this axis0
2020APT29211409.3% to 37.1%0
2020Carbanak+FIN72917914.5% to 87.2%0
2022Wizard Spider + Sandworm3010918.3% to 98.2%1
2023Turla · this page2914323.8% to 99.3%5
2024Enterprise 20241910011% to 80%0
2025Enterprise 20251110720.6% to 84.1%0

The documents this page rests on

The results themselves are at MITRE’s evaluations site, and this round sits in the enterprise results index. The adversary the attack was modelled on has its own ATT&CK page, and the machine-readable form of the evaluations, which is what was actually parsed, is the ATT&CK Evaluations Library. Every technique id in the substeps table opens the ATT&CK page for that technique. Where this page and MITRE disagree, MITRE is right and we would like to know.

Ask us to read a record

There is a number in a slide deck about a product you are about to buy, or already run. It came from a public record, and the record almost always says more than the number does — sometimes something that changes the decision, sometimes nothing at all.

Send us the claim and the record behind it, and we will do to it what this page does to this one: read the second axis, write down what the record can and cannot answer, and publish the result whichever way it goes, including the dull way where the number holds up and the page says so. We read everything and we answer, including when the answer is no.

9592 Solutions UG (haftungsbeschränkt), Fährstr. 217, 40221 Düsseldorf, Germany is the controller for what you send here. Your address and your message are used to answer you and to work out whether we take the request on, under Art. 6(1)(b) and Art. 6(1)(f) GDPR. They go to nobody else and they are not used for advertising. Write to christo@9592.tech for a copy or a deletion at any time. The longer version is on the privacy page.

9592 Solutions UG (haftungsbeschränkt), Düsseldorf · Privacy · youtube.com/@m37channel