Watch Your Language: Detecting Multi-Lingual Language-Confusion Phishing

Detects phishing messages that deliberately mix writing systems/scripts/homoglyphs in the Subject line as ONE corroborating signal within a broader, multi-signal phishing pattern. Hardened against false positives: brand/product tokens (e.g. Microsoft, PayPal) are stripped before the Latin-script check so a lone brand mention in non-Latin correspondence doesn't count as mixed-script; a weighted score (script-mixing=2, homoglyph=3, high-confidence phishing vocabulary=2, threshold=6) means no two of the three signal categories can reach the threshold alone — all three must co-occur before the rule fires.