Papers
Topics
Authors
Recent
Search
2000 character limit reached

RoBERTa-Augmented Synthesis for Detecting Malicious API Requests

Published 18 May 2024 in cs.CR | (2405.11258v3)

Abstract: Web applications and APIs face constant threats from malicious actors seeking to exploit vulnerabilities for illicit gains. To defend against these threats, it is essential to have anomaly detection systems that can identify a variety of malicious behaviors. However, a significant challenge in this area is the limited availability of training data. Existing datasets often do not provide sufficient coverage of the diverse API structures, parameter formats, and usage patterns encountered in real-world scenarios. As a result, models trained on these datasets often struggle to generalize and may fail to detect less common or emerging attack vectors. To enhance detection accuracy and robustness, it is crucial to access larger and more representative datasets that capture the true variability of API traffic. To address this, we introduce a GAN-inspired learning framework that extends limited API traffic datasets through targeted, domain-aware synthesis. Drawing on techniques from NLP, our approach leverages Transformer-based architectures, particularly RoBERTa, to enhance the contextual representation of API requests and generate realistic synthetic samples aligned with security-specific semantics. We evaluate our framework on two benchmark datasets, CSIC 2010 and ATRDF 2023, and compare it with a previous data augmentation technique to assess the importance of domain-specific synthesis. In addition, we apply our augmented data to various anomaly detection models to evaluate its impact on classification performance. Our method achieves up to a 4.94% increase in F1 score on CSIC 2010 and up to 21.10% on ATRDF 2023. The source codes of this work are available at https://github.com/ArielCyber/GAN-API.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (38)
  1. R. Sun, Q. Wang, and L. Guo, “Research Towards Key Issues of API Security,” in China Cyber Security Annual Conference, 2021, pp. 179–192.
  2. F. Hussain, B. Noye, and S. Sharieh, “Current state of API security and machine learning,” IEEE Technology Policy and Ethics, vol. 4, no. 2, pp. 1–5, 2019.
  3. M. Idris, I. Syarif, and I. Winarno, “Web Application Security Education Platform Based on OWASP API Security Project,” International Journal of Engineering Technology, pp. 246–261, 2022.
  4. Salt.security. (2021) API Security Trends. [Accessed: 23-November-2021]. [Online]. Available: https://salt.security/api-securitytrends
  5. K. Chang, N. Zhao, and L. Kou, “A Survey on Malware Detection based on API Calls,” in 2022 9th International Conference on Dependable Systems and Their Applications (DSA), 2022, pp. 464–471.
  6. Y. Yu and N. Bian, “An intrusion detection method using few-shot learning,” IEEE Access, vol. 8, pp. 49 730–49 740, 2020.
  7. J. Bugeja and J. A. Persson, “A data-centric anomaly-based detection system for interactive machine learning setups,” in 18th International Conference on Web Information Systems and Technologies-WEBIST.   SciTePress, 2022.
  8. K. Parnow, Z. Li, and H. Zhao, “Grammatical Error Correction as GAN-like Sequence Labeling,” in Association for Computational Linguistics: ACL-IJCNLP, 2021, pp. 3284–3290.
  9. Z. Li, S. Chen, H. Dai, D. Xu, C.-K. Chu, and B. Xiao, “Abnormal Traffic Detection: Traffic Feature Extraction and DAE-GAN With Efficient Data Augmentation,” IEEE Transactions on Reliability, 2022.
  10. L. Sun, Y. Zhou, Y. Wang, C. Zhu, and W. Zhang, “The effective methods for intrusion detection with limited network attack data: Multi-task learning and oversampling,” IEEE Access, vol. 8, pp. 185 384–185 398, 2020.
  11. N.-T. Tran, T.-A. Bui, and N.-M. Cheung, “Dist-gan: An improved gan using distance constraints,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 370–385.
  12. Z. Du, L. Gao, and X. Li, “A new contrastive GAN with data augmentation for surface defect recognition under limited data,” IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1–13, 2022.
  13. Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” arXiv preprint arXiv:1907.11692, 2019.
  14. M. Gniewkowski, H. Maciejewski, T. R. Surmacz, and W. Walentynowicz, “HTTP2vec: Embedding of HTTP Requests for Detection of Anomalous Traffic,” arXiv preprint arXiv:2108.01763, 2021.
  15. D. Chen, Q. Yan, C. Wu, and J. Zhao, “SQL Injection Attack Detection and Prevention Techniques Using Deep Learning,” in Journal of Physics: Conference Series, vol. 1757, 2021, p. 012055.
  16. K. Zhang, “A machine learning-based approach to identify SQL Injection vulnerabilities,” in 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE).   IEEE, 2019, pp. 1286–1288.
  17. Y. Pan et al., “Detecting web attacks with end-to-end deep learning,” Journal of Internet Services and Applications, vol. 10, no. 1, pp. 1–22, 2019.
  18. X. Xie et al., “SQL Injection detection for web applications based on Elastic-Pooling CNN,” IEEE Access, vol. 7, pp. 151 475–151 481, 2019.
  19. A. A. Ashlam, A. Badii, and F. Stahl, “A novel approach exploiting machine learning to detect sqli attacks,” in ICASET, 2022, pp. 513–517.
  20. J. Li et al., “Web application attack detection based on attention and gated convolution networks,” IEEE Access, vol. 8, pp. 20 717–20 724, 2019.
  21. S. Huang, Y. Liu, C. Fung, W. An, R. He, Y. Zhao, H. Yang, and Z. Luan, “A gated few-shot learning model for anomaly detection,” in 2020 International Conference on Information Networking (ICOIN), 2020, pp. 505–509.
  22. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017.
  23. T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz et al., “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2020, pp. 38–45.
  24. C. Wang, K. Cho, and J. Gu, “Neural machine translation with byte-level subwords,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 05, 2020, pp. 9154–9160.
  25. H. T. Kesgin and M. F. Amasyali, “Iterative mask filling: An effective text augmentation method using masked language modeling,” in International Conference on Advanced Engineering, Technology and Applications.   Springer, 2023, pp. 450–463.
  26. P. J. Rousseeuw and M. Hubert, “Robust statistics for outlier detection,” Wiley Interdisciplinary Reviews: Data mining and knowledge discovery, vol. 1, no. 1, pp. 73–79, 2011.
  27. A.-M. Simundic et al., “Confidence interval,” Biochemia Medica, vol. 18, no. 2, pp. 154–161, 2008.
  28. D. Bear and P. Cook, “Fine-tuning Sentence-RoBERTa to Construct Word Embeddings for Low-resource Languages from Bilingual Dictionaries,” in Proceedings of the Workshop on Natural Language Processing for Indigenous Languages of the Americas (AmericasNLP), 2023, pp. 47–57.
  29. Y. Ho and S. Wookey, “The real-world-weight cross-entropy loss function: Modeling the costs of mislabeling,” IEEE access, vol. 8, pp. 4806–4813, 2019.
  30. A. Manolache, F. Brad, and E. Burceanu, “Date: Detecting anomalies in text via self-supervision of transformers,” arXiv preprint arXiv:2104.05591, 2021.
  31. T. C. Rajapakse, “Simple transformers,” https://github.com/ThilinaRajapakse/simpletransformers, 2019.
  32. A. Podolskiy, D. Lipin, A. Bout, E. Artemova, and I. Piontkovskaya, “Revisiting mahalanobis distance for transformer-based out-of-domain detection,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 15, 2021, pp. 13 675–13 682.
  33. S. Shukla, S. Sarthak, and K. V. Arya, “Noobs at semeval-2021 task 4: Masked language modeling for abstract answer prediction,” in Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), 2021, pp. 805–809.
  34. C. T. Giménez, A. P. Villegas, and G. Á. Marañón, “HTTP data set CSIC 2010,” CSIC, vol. 64, 2010.
  35. S. Emanet, G. K. Baydogmus, and O. Demir, “An ensemble learning based IDS using Voting rule: VEL-IDS,” PeerJ Computer Science, vol. 9, p. e1553, 2023.
  36. C. R. Pardomuan, A. Kurniawan, M. Y. Darus, M. A. Mohd Ariffin, and Y. Muliono, “Server-Side Cross-Site Scripting Detection Powered by HTML Semantic Parsing Inspired by XSS Auditor.” Pertanika Journal of Science & Technology, vol. 31, no. 3, 2023.
  37. S. Lavian, R. Dubin, and A. Dvir, “The API Traffic Research Dataset Framework (ATRDF),” 2023, https://github.com/ArielCyber/Cisco_Ariel_Uni_API_security_challenge.
  38. K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 2002, pp. 311–318.
Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.