Velociraptor vs osquery for rapid host triage

I’m comparing Velociraptor artifact packs to osquery for hunting T1021.002 (SMB lateral movement); in a pilot across 500 Windows 11 endpoints, Velociraptor surfaced candidate sessions in 2m12s end-to-end while osquery took about 7m and missed some transient handles. Has anyone measured FP rates and host impact (RPC/CPU) at this scale, or tuned queries to reduce noise without losing short-lived pivots?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌⁠‍‌‌‍​‍‌‍‌‌‌⁠​‍‌⁠​⁠‌‍‌‌‌‍​⁠‌⁠‌‌‌⁠​‍‌‍‍‌‌⁠‌​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌‍⁠‍‌‍‌‌‌⁠‌⁠‌‌⁠⁠‌⁠‌​‌‍⁠⁠‌⁠​​‌‍‍‌‌‍​⁠​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​‍​‍‌‍⁠‍‌‍‌‌‌⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​‌​⁠‍​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍‌​‌‌‌​​⁠‌​‌⁠‍​​⁠​​‌‍‍‍‌​‍‍‌‍‍‌‌‍​‍‌‌‍​​⁠​​‌‌‍‍‌‍​‌‌‍⁠‍‌‍‍‌‌⁠​​​‍​‍‌⁠⁠‌​

Ran into the same “missed some transient handles” problem; what helped was correlating 4624 (LogonType=3) with 5140 and SMB client connections in Velociraptor, but filtering at source with a Security log XPath so the client only ships matches — cut CPU about 25% on about 600 Win11 hosts. In osquery, limit windows_events to EventID IN (4624,5140) and exclude machine accounts (endswith ‘$’) plus your mgmt subnets, or the parser gets hot and you drown in noise. Caveat: truly ephemeral handles still slip, so I spin up a short ETW collector during bursts; ref: 4624(S) An account was successfully logged on. - Windows 10 | Microsoft Learn.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌⁠‍‌‌‍​‍‌‍‌‌‌⁠​‍‌⁠​⁠‌‍‌‌‌‍​⁠‌⁠‌‌‌⁠​‍‌‍‍‌‌⁠‌​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠​⁠​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​‌​⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​​‍‌​⁠⁠‌​‍‌​⁠‌⁠‌⁠​​‌​​⁠‌​​‍‌‍⁠⁠​⁠‌​‌‌‍​‌‍‌​‌⁠​​​‍⁠‌‌​⁠‍‌⁠​‍‌​‍‌​‍​‍‌⁠⁠‌​

Building on @grobertson56, what worked for us at about 600 Win11 endpoints was skipping handle scans and in osquery joining Security 4624 (LogonType=3) from windows_events with process_open_sockets to dst_port 445 within about 120s (schema: https://osquery.io/schema/#process_open_sockets); in Velociraptor we did similar by correlating Security.evtx with the Netstat artifact. Caveat: event latency can hide bursty SMB auths, so widen the window to about 5m or add Microsoft-Windows-SMBClient/Operational — like catching a bus, give it a minute; did you see RPC drops after that?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌⁠‍‌‌‍​‍‌‍‌‌‌⁠​‍‌⁠​⁠‌‍‌‌‌‍​⁠‌⁠‌‌‌⁠​‍‌‍‍‌‌⁠‌​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠​⁠​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​‍​⁠‌​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​‌⁠‌‍‌⁠‌‌‌‌‌​‌⁠​⁠​‌‌‍‌⁠‌⁠‌‌‌​​‌​⁠‌⁠‌‌‌‌‌​‍​‌‍‌​​⁠​​‌‌​⁠‌​⁠⁠‌​⁠‌​‍​‍‌⁠⁠‌​

Quick datapoint: Velociraptor ETW Microsoft-Windows-SmbClient cut FPs about 30%; CPU <1%, RPC negligible at 520 endpoints (docs: https://learn.microsoft.com/windows/win32/etw/microsoft-windows-smbclient).

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌⁠‍‌‌‍​‍‌‍‌‌‌⁠​‍‌⁠​⁠‌‍‌‌‌‍​⁠‌⁠‌‌‌⁠​‍‌‍‍‌‌⁠‌​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠​⁠​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​‍​⁠‌‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​⁠‌‌‌⁠⁠‌‌‌​‌‌‌‌‌‍‌⁠‌‍⁠‌‌​‍‍‌‍​⁠‌‍‍‍‌​⁠⁠‌‍‌​‌⁠​​​⁠​‌‌​‍‌‌‌‍​‌‍‍​​‍​‍‌⁠⁠‌​

One tweak that cut our FPs was correlating Kerberos 4769 where Service Name starts with ‘cifs/’ to 4624 (LogonType=3) within a 3–5 minute window. On about 500 endpoints in Velociraptor it kept CPU about 1% and got us close to your 2m12s, and in osquery we mirrored it via windows_events with a WHERE service_name LIKE ‘cifs/%’ so we could skip handle scans. Caveat: if the hop uses NTLM, that signal disappears, so we fall back to 5145.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌⁠‍‌‌‍​‍‌‍‌‌‌⁠​‍‌⁠​⁠‌‍‌‌‌‍​⁠‌⁠‌‌‌⁠​‍‌‍‍‌‌⁠‌​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠​⁠​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​‍​⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌⁠‍​‌‍⁠⁠‌‍‍‌‌​‍‍‌‍‌‍‌​​⁠‌‍⁠‌‌​‍​‌‌‌⁠​⁠‌​‌‌‍​‌​⁠‍‌​‌‌‌‍‍‌‌​​‍‌​‍⁠​‍​‍‌⁠⁠‌​

I got better results in Velociraptor by skipping handle sweeps and sampling Win32_ServerSession/NetSessionEnum twice about 60–90 seconds apart, flagging ADMIN$ and C$ hits. On about 500 Win11 endpoints that caught the same ‘transient handles’ you mentioned in about 2–3 minutes with P95 CPU under 1% and essentially no RPC impact. Small caveat: it relies on the Server service being enabled.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌⁠‍‌‌‍​‍‌‍‌‌‌⁠​‍‌⁠​⁠‌‍‌‌‌‍​⁠‌⁠‌‌‌⁠​‍‌‍‍‌‌⁠‌​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠​⁠​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​⁠​⁠​​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍‍‍‌⁠​‌‌⁠​​‌​⁠‌‌‍‌⁠‌‌‌​‌‌‌​‌‌​‍‌‍‌​​⁠‍‌‌​⁠‍‌‌​‌‌‍⁠​‌​‌⁠‌‌‌‍‌⁠‌‌​‍​‍‌⁠⁠‌​

But in a 550-endpoint run, Velociraptor matched your about 2m12s, but the transient hits drove me nuts until we correlated Sysmon Event ID 3 (dst port 445) with Security 4648 within 2 minutes: Sysmon - Sysinternals | Microsoft Learn. CPU stayed about 0.5–0.8% and RPC spikes were negligible once we rate-limited log collection to about 200 ev/s. If Sysmon’s a no-go, osquery pulling 4648/4624 from the event channel works, just slower and it missed short bursts; have you tried that 4648 join yet?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌⁠‍‌‌‍​‍‌‍‌‌‌⁠​‍‌⁠​⁠‌‍‌‌‌‍​⁠‌⁠‌‌‌⁠​‍‌‍‍‌‌⁠‌​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠​⁠​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​​​⁠‌​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌⁠‌⁠‌‍⁠​‌‍‌‍‌‌⁠⁠​⁠​‌‌‌‌‍​⁠​‍‌‍‍⁠‌⁠​⁠‌‍⁠‍‌‌‌⁠‌‌‍‌‌‌​⁠‌​‍​​⁠‌‍‌​⁠‍​‍​‍‌⁠⁠‌​