ReversingLabs

What RL Found Before Anthropic’s Midnight Blizzard Report

Anthropic’s September report revealed that a Russia-linked threat actor was rebuilding flagged software implants with AI, making it hard to detect. 

However, ReversingLabs (RL) flagged the rebuilt malware two months before Anthropic did. And our data on the same malware shows that, while impressive, such AI powered tactics are not decisive. Instead, threat context behaviors and history are more critical than ever in spotting malware and improving the outcomes for targeted organizations. 

Here’s the background: Anthropic’s Threat Intelligence Report, released in September 2026 describes GTG-20006, a Russian-speaking espionage operator whose attribution Anthropic calls “consistent with public reporting linking the actor to Midnight Blizzard.” Anthropic’s report said, “when their implants were flagged by security products, the actor used Claude to systematically identify, modify and redeploy the detected artifacts.”

That might seem like it’s confirming what doomsayers have long predicted, that large language model AI will throw a wrench into threat detection — in this case, by completely re-writing malware at machine speed. 

But the story is not that simple. RL ran the case’s two malware hashes against our own telemetry. What we found is that, while AI may have changed how fast an operator can rebuild a file, it did not change the key identifiers that threat  analysts look for to determine whether software is suspicious, malicious, or legitimate. In this case the answers came from data that no rebuild loop erases: Every hash, rule and query is in the full technical analysis.

Key findings

  • Both files were already in RL telemetry. RL first observed the Go backdoor on July 6 and the PowerShell stager on July 10, about two months before the report, and classified both as malicious on August 1.
  • Five builds shared one behavior profile. Five builds of the stage malware reached RL between July 9 and August 3 with five different hashes. A hunt written on the operator’s beacon labels returned 7 malicious files and 0 goodware.
  • Reputation lagged. Thirteen of the 14 servers Anthropic lists looked low risk to third-party reputation sources. The one RL tied to malware had a single flag before the report.
  • Two hashes became a map. RL found 7 more files, 8 URLs and 2 links between the cluster’s domains and servers, none of them in Anthropic’s appendix.

What AI Changed

In a word: Speed. An operator who rebuilds an implant by hand does it a few times. An operator who hands the job to a model does it every time a scanner flags the file. Anthropic describes that loop, and it makes every published hash go stale faster.

RL’s data shows the same pattern, though it cannot show whether AI produced it. Five stager builds with five hashes reached RL in 25 days, and fuzzy hashing splits them into two code variants. For a team that defends on hashes, each build restarts the clock.

Anthropic warns that AI “threatens to quickly and easily subvert defenders’ ability to impose costs on adversaries via static detections alone.” 

I agree — as far as static detections go. But those are hardly the only tools available to defenders. 

What AI Did Not Change

Hand an analyst a fresh indicator and four questions follow in the same order: 

  1. Have we seen this before? 
  2. What does it do? 
  3. What is it connected to? And… 
  4. Where else does it appear? 

An AI assistant working on the attacker’s side changes the answer to none of these, because they are focused on what an attack does rather than what it is.

Question

What RL data answered

Seen before?

Yes. First observed July 6 and July 10; classified malicious August 1.

What does it do?

Runs hidden code, steals browser data, persists and calls home, with the same profile in every build.

What is it connected to?

A staging server that received the stager’s beacons, hosted two more stagers and was the target of two lure shortcuts.

Where else does it appear?

In 7 more files, 8 URLs and 2 domain-to-server links.

Figure 1. How two published hashes became a campaign map: one RL pivot per row, and a strip at the bottom counting what Anthropic’s appendix does not contain.

The answers to each of those questions are critical. But it is the answer to the second question, “What does it do?” that is the heart of the matter. In the case of the Anthropic malware, the stager hides its window, decodes and runs hidden code, collects browser history with public PowerSploit/Empire code and reports to a staging server.

Likewise, the backdoor installs under a fake Windows service name, restarts with the machine, injects into another process and reads browser and Teams data.

Those behaviors have had MITRE ATT&CK entries for years. AI did not give the malware new ways to work or new, unseen behaviors. As a result, a hunt built on well-identified malicious behaviors set off red flags and returned malicious files only.

Where Reputation Fell Short

The staging server stood out because of the files linked to it. It received the stager’s beacons, hosted two more stagers and was the target of both lure shortcuts. Its reputation lagged well behind the file evidence RL had that indicated malicious activity. And that’s a big problem for defender organizations. Of the six sources that flag the staging server today, one did so before the report.

The lure shortcuts show the same detection gap from the file side. They contain no code, and no antivirus engine flags them. Even RL’s consensus-based cloud verdict still calls them goodware. Their malicious nature is based entirely on where they point. Of the 61 domains in Anthropic’s appendix that RL now classifies as malicious, 48 were flagged as malicious in RL’s telemetry before publication, yet the median domain is still flagged by only 2 out of about 47 reputation sources. Presence is not a verdict. It is the history that makes one possible.

What Stopped Working

The signals that failed were the ones that describe how a file was built. The Go backdoor’s import hash matches 12,000 goodware and 6,900 malicious files in RL telemetry over the past year, because Go programs built with the same toolchain look alike regardless of what they do. Also, the Go backdoor’s install filename collides with 14 unrelated malware samples from families that share nothing with this cluster. 

Likewise, a search on its hash returns about 19,500 similar files across dozens of families. Build artifacts, names and hashes are cheap to change using AI, so their failure here is the expected result.

What This Means for Security Operations Leaders

Here’s what you need to know:

  • Retro-hunt every vendor report the day it lands. Both seed files here had been in RL’s corpus for two months. The first question only has a useful answer when your data reaches back far enough.
  • Invest in hunting on behavior and operator choices. The beacon labels found five stagers the published hashes could not.
  • Treat reputation as one input. Enrich servers with the files and domains linked to them before deciding they are harmless.
  • Plan for hash indicators to age faster. Build detection content around capabilities, and measure how long hash-based blocks stay useful in your environment.

At RL, our Spectra Intelligence telemetry provided the critical threat history data, while the behavioral hunt was fueled by data from  YARA retro-hunts using Spectra Analyze. 

Your teams can run the same queries and rules from the technical post on their own data.

In the end, Anthropic’s Threat Intelligence Report documents a clear trend: AI powered threat acceleration. But RL’s telemetry documents why that, alone, isn’t enough to avoid detection. Instead, defenders and threat analysts need to stay focus on what stays fixed. In the end, five AI-generated stager builds reached RL, and none of them shed the need to run unseen, collect data and call home. 

Where are things headed? We assess with moderate confidence that AI-assisted operators will keep shortening the useful life of hash indicators. We judge it highly likely that historical telemetry and the relationships between files and infrastructure will gain value as that happens. The analyst’s questions will stay the same. The depth of the data behind the answers will decide who wins.

See the press release: ReversingLabs Wins 2026 CyberSecurity Breakthrough Award for Threat Detection Solution Provider of the Year

Leave a Reply