Texts in phishing emails are more elaborate and do not contain explicit triggers, unlike, for example, spam texts. The latter use vocabulary from very specific areas, such as pharmaceuticals, which are easy to detect. A Viagra advert is a spam classic. Phishing texts vary greatly and change over time, which makes them difficult to be captured by signature methods alone and do not allow detection solutions to make a verdict based solely on the text whether the email is malicious or not.
The second classifier uses a model to exactly analyze incoming messages and detect phishing phrases.
During the training period, the model analyzes many examples of phishing emails: it splits them into separate phrases and assigns weight to each phrase depending on its potential for phishing activity, or how common it is among phishing communications. For example, the phrase “Best regards” will carry a low weight because it has no signs of phishing and is often used in legitimate messages. The phrases that call for payment, following a link and entering data will have more weight, because they are specifically found in phishing emails.
As a result of this weighting, the model gains a whole category of words and phrases which will then be considered as suspicious if found in a message. This category is passed to the classifier, which is located on the client’s device (see image three). The classifier uses this category as a reference and based on this, it can decide whether the messages arriving in the client’s mailbox contain phishing context.
This method is conceptually simple, but highly effective, because, firstly, the algorithm of the model is transparent, or interpretable – it is easy to understand by the weight why the model made its verdict. And secondly, it quickly learns and quickly updates. This helps the model maintain its performance and quality of detection due to the evolution of phishing texts over time.
What is the Result?
By combining the classifiers that inspect the content of an email and its header metadata, we achieved a new technology that can detect and confidently block phishing emails in real time. The verdict of only one of the two classifiers is not enough to define if a message is phishing. It is necessary for the verdicts to coincide – this allows the technology to more accurately identify malicious messages and minimize the likelihood of a false positive.
What is the advantage of this solution? Firstly, this allows us to be much more proactive. The technology is able to detect phishing techniques that have not been seen before. This distinguishes the solution from popular methods based on the signature approach, when you first need to see the sample in order to block it in the future.
Secondly, it is a fully automatic solution. The technology independently learns from new collections of statistical data, which allows the speed of response and detection of malicious emails to increase.
Thirdly, it can detect non-trivial patterns. Thanks to the model architecture and the large amount of data available for training, it is able to extract complex patterns. For example, this includes strange sequences of characters or registers in headers that cannot be found manually.
We continue to improve the technology and plan to add other classifiers that will analyze more message parameters. In its current form, as part of the solutions for protecting mail servers and the Microsoft Office 365 application, the technology has already been used to help increase the detection rate of the most sophisticated phishing emails.
Explore Kaspersky Security Solutions for Enterprise to predict, prevent, detect and respond to cyber-attacks.
