Hi Google SecOps Community,
i would like to confirm whether YARA-L supports Levenshtein distance (edit distance) or any equivalent fuzzy string-matching functionality.
Use Case:
We have a detection that identifies potential data exfiltration through Microsoft 365 Exchange emails by comparing the username portion of the sender's corporate email address with the username portion of the recipient's personal email address.
The objective is to identify sender and recipient usernames that are identical or differ by fewer than three character edits.
For example:
-
[removed by moderator]→[removed by moderator](distance = 1) -
[removed by moderator]→[removed by moderator](distance = 2)
Questions:
-
Does Google SecOps YARA-L 2.0 provide a built-in function for calculating Levenshtein distance between two strings?
-
If Levenshtein distance is not supported natively, is there another supported YARA-L function or approach for fuzzy string matching?
-
Can this comparison be implemented directly within a YARA-L detection rule using extracted sender and recipient usernames?
-
If native support is unavailable, what is Google's recommended approach for implementing this detection? Would upstream enrichment be required?
-
Are there any existing Google SecOps detection rules or examples that implement similar username-similarity logic?
Any guidance, documentation, or working YARA-L examples would be greatly appreciated.
Thanks!



