Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Classification-by-Retrieval Framework for Few-Shot Anomaly Detection to Detect API Injection Attacks

Published 18 May 2024 in cs.CR | (2405.11247v2)

Abstract: Application Programming Interface (API) Injection attacks refer to the unauthorized or malicious use of APIs, which are often exploited to gain access to sensitive data or manipulate online systems for illicit purposes. Identifying actors that deceitfully utilize an API poses a demanding problem. Although there have been notable advancements and contributions in the field of API security, there remains a significant challenge when dealing with attackers who use novel approaches that don't match the well-known payloads commonly seen in attacks. Also, attackers may exploit standard functionalities unconventionally and with objectives surpassing their intended boundaries. Thus, API security needs to be more sophisticated and dynamic than ever, with advanced computational intelligence methods, such as machine learning models that can quickly identify and respond to abnormal behavior. In response to these challenges, we propose a novel unsupervised few-shot anomaly detection framework composed of two main parts: First, we train a dedicated generic LLM for API based on FastText embedding. Next, we use Approximate Nearest Neighbor search in a classification-by-retrieval approach. Our framework allows for training a fast, lightweight classification model using only a few examples of normal API requests. We evaluated the performance of our framework using the CSIC 2010 and ATRDF 2023 datasets. The results demonstrate that our framework improves API attack detection accuracy compared to the state-of-the-art (SOTA) unsupervised anomaly detection baselines.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (75)
  1. S. Balsari, A. Fortenko, J. A. Blaya, A. Gropper, M. Jayaram, R. Matthan, R. Sahasranam, M. Shankar, S. N. Sarbadhikari, B. E. Bierer et al., “Reimagining Health Data Exchange: An application programming interface–enabled roadmap for India,” Journal of Medical Internet Research, vol. 20, no. 7, p. e10725, 2018.
  2. J. Ofoeda, R. Boateng, and J. Effah, “Application programming interface (API) research: A review of the past to inform the future,” IJEIS, vol. 15, no. 3, pp. 76–95, 2019.
  3. A. Mendoza and G. Gu, “Mobile application web API reconnaissance: Web-to-mobile inconsistencies & vulnerabilities,” in SP, 2018, pp. 756–769.
  4. C. Benzaid and T. Taleb, “ZSM security: Threat surface and best practices,” IEEE Network, vol. 34, no. 3, pp. 124–133, 2020.
  5. IBM, “Innovation in the API economy: Building winning experiences and new capabilities to compete,” 2016, https://www.ibm.com/downloads/cas/OXV3LYLO.
  6. R. Sun, Q. Wang, and L. Guo, “Research Towards Key Issues of API Security,” in CNCERT, Beijing, China, July 20–21, 2021, pp. 179–192.
  7. M. Coyne, “Zoom’s Big Security Problems Summarized,” https://www.forbes.com/sites/marleycoyne/2020/04/03/zooms-big-security-problems-summarized/?sh=46fc370f4641, 2020, forbes.
  8. “Capital One data breach: Arrest after details of 106m people stolen,” https://www.bbc.com/news/world-us-canada-49159859, 2019, bBC.
  9. J. Greig, “Hilton denies hack after data from 3.7 million Honors customers offered for sale,” https://therecord.media/hilton-denies-hack-after-data-from-3-7-million-honors-customer-offered-for-sale/, 2023, the Record.
  10. “Equifax Says Cyberattack May Have Affected 143 Million in the U.S.” https://www.nytimes.com/2017/09/07/business/equifax-cyberattack.html, 2017, the New York Times.
  11. D. Fett, R. Küsters, and G. Schmitz, “A comprehensive formal security analysis of OAuth 2.0,” in SIGSAC CCCS, 2016, pp. 1204–1215.
  12. A. Chan, A. Kharkar, R. Z. Moghaddam, Y. Mohylevskyy, A. Helyar, E. Kamal, M. Elkamhawy, and N. Sundaresan, “Transformer-based Vulnerability Detection in Code at EditTime: Zero-shot, Few-shot, or Fine-tuning?” arXiv preprint arXiv:2306.01754, 2023.
  13. G. Baye, F. Hussain, A. Oracevic, R. Hussain, and S. A. Kazmi, “API security in large enterprises: Leveraging machine learning for anomaly detection,” in ISNCC.   IEEE, 2021, pp. 1–6.
  14. E. Harlicaj et al., “Anomaly detection of web-based attacks in microservices,” 2021.
  15. F. Shen, Y. Mu, Y. Yang, W. Liu, L. Liu, J. Song, and H. T. Shen, “Classification by retrieval: Binarizing data and classifiers,” in ACM SIGIR, 2017, pp. 595–604.
  16. W. Shi, J. Michael, S. Gururangan, and L. Zettlemoyer, “Nearest neighbor zero-shot inference,” arXiv preprint arXiv:2205.13792, 2022.
  17. A. M. Qamar, E. Gaussier, J.-P. Chevallet, and J. H. Lim, “Similarity learning for nearest neighbor classification,” in ICDM.   IEEE, 2008, pp. 983–988.
  18. J. J. Valero-Mas, A. J. Gallego, P. Alonso-Jiménez, and X. Serra, “Multilabel prototype generation for data reduction in k-nearest neighbour classification,” Pattern Recognition, vol. 135, p. 109190, 2023.
  19. A. S. Reddy and B. Rudra, “Evaluation of Recurrent Neural Networks for Detecting Injections in API Requests,” in CCWC, 2021, pp. 0936–0941.
  20. J. Ombagi, “Time-Based Blind SQL Injection via HTTP Headers: Fuzzing and Exploitation.” 2017.
  21. I. Jemal, M. A. Haddar, O. Cheikhrouhou, and A. Mahfoudhi, “Performance evaluation of Convolutional Neural Network for web security,” Computer Communications, vol. 175, pp. 58–67, 2021.
  22. M. Gniewkowski, H. Maciejewski, T. R. Surmacz, and W. Walentynowicz, “HTTP2vec: Embedding of HTTP Requests for Detection of Anomalous Traffic,” ArXiv, vol. abs/2108.01763, 2021.
  23. I. Jemal, M. A. Haddar, O. Cheikhrouhou, and A. Mahfoudhi, “M-CNN: a new hybrid deep learning model for web security,” in AICCSA, 2020, pp. 1–7.
  24. Q. Niu and X. Li, “A high-performance web attack detection method based on CNN-GRU model,” in ITNEC, vol. 1, 2020, pp. 804–808.
  25. L. Yu, L. Chen, J. Dong, M. Li, L. Liu, B. Zhao, and C. Zhang, “Detecting malicious web requests using an enhanced textCNN,” in COMPSAC.   IEEE, 2020, pp. 768–777.
  26. A. Moradi Vartouni, S. Mehralian, M. Teshnehlab, and S. Sedighian Kashi, “Auto-Encoder LSTM Methods for Anomaly-Based Web Application Firewall,” International Journal of Information and Communication Technology Research, vol. 11, no. 3, pp. 49–56, 2019.
  27. A. Moradi Vartouni, M. Shokri, and M. Teshnehlab, “Auto-threshold deep SVDD for anomaly-based web application firewall,” 2021.
  28. F. A. Research, “fastText library for efficient learning of word representations and sentence classification,” https://github.com/facebookresearch/fastText/, 2017, online; accessed 02-December-2017.
  29. B. Bansal and S. Srivastava, “Sentiment classification of online consumer reviews using word vector representations,” Procedia Computer Science, vol. 132, pp. 1147–1153, 2018.
  30. S. TOPRAK and A. G. YAVUZ, “Web Application Firewall Based on Anomaly Detection Using Deep Learning,” Acta Infologica, vol. 6, no. 2, pp. 219–244, 2022.
  31. L. Xiao, S. Matsumoto, T. Ishikawa, and K. Sakurai, “SQL Injection Attack Detection Method Using Expectation Criterion,” in 2016 Fourth International Symposium on Computing and Networking (CANDAR).   IEEE, 2016, pp. 649–654.
  32. Y. E. Seyyar, A. G. Yavuz, and H. M. Ünver, “An attack detection framework based on BERT and deep learning,” IEEE Access, vol. 10, pp. 68 633–68 644, 2022.
  33. E. Chávez, G. Navarro, R. Baeza-Yates, and J. L. Marroquín, “Searching in metric spaces,” ACM Computing Surveys (CSUR), vol. 33, no. 3, pp. 273–321, 2001.
  34. A. Ponomarenko, N. Avrelin, B. Naidan, and L. Boytsov, “Comparative analysis of data structures for approximate nearest neighbor search,” in DATA ANALYTICS, 2014, pp. 125–130.
  35. P. Indyk and R. Motwani, “Approximate nearest neighbors: towards removing the curse of dimensionality,” in ACM Symposium on Theory of Computing, 1998, pp. 604–613.
  36. K. Hajebi, Y. Abbasi-Yadkori, H. Shahbazi, and H. Zhang, “Fast approximate nearest-neighbor search with k-nearest neighbor graph,” in AI, 2011.
  37. R. Battle and E. Benson, “Bridging the semantic Web and Web 2.0 with representational state transfer (REST),” Journal of Web Semantics, vol. 6, no. 1, pp. 61–69, 2008.
  38. W. J. Buchanan, S. Helme, and A. Woodward, “Analysis of the adoption of security headers in HTTP,” IET Information Security, vol. 12, no. 2, pp. 118–126, 2018.
  39. A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov, “Bag of Tricks for Efficient Text Classification,” arXiv preprint arXiv:1607.01759, 2016.
  40. B. Li and L. Han, “Distance weighted cosine similarity measure for text classification,” in IDEAL, China, Oct., 2013, pp. 611–618.
  41. E. Techapanurak, M. Suganuma, and T. Okatani, “Hyperparameter-free out-of-distribution detection using cosine similarity,” in ACCV, 2020.
  42. Y.-C. Hsu, Y. Shen, H. Jin, and Z. Kira, “Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data,” in CVF, 2020, pp. 10 951–10 960.
  43. P. Xia, L. Zhang, and F. Li, “Learning similarity with cosine similarity ensemble,” Information Sciences, vol. 307, pp. 39–52, 2015.
  44. B. Naidan, L. Boytsov, Y. Malkov, and D. Novak, “Non-Metric Space Library (NMSLIB): An efficient similarity search library and a toolkit for evaluation of k-NN methods for generic non-metric spaces,” https://github.com/nmslib/nmslib, 2014, online; accessed 14-July-2014.
  45. J. E. Stone, K. S. Griffin, J. Amstutz, D. E. DeMarle, W. R. Sherman, and J. Günther, “ANARI: A 3-D Rendering API Standard,” Computing in Science & Engineering, vol. 24, no. 2, pp. 7–18, 2022.
  46. C. Torrano-Gimenez, H. T. Nguyen, G. Alvarez, S. Petrović, and K. Franke, “Applying feature selection to payload-based web application firewalls,” in IWSCN, 2011, pp. 75–81.
  47. B. Ware, “Analyzing Web Traffic ECML/PKDD 2007 Discovery Challenge,” 2007, https://www.lirmm.fr/pkdd2007-challenge/comites.html.
  48. UNB, “A collaborative project between the Communications Security Establishment (CSE) & the Canadian Institute for Cybersecurity (CIC),” 2018, https://www.unb.ca/cic/datasets/ids-2018.html.
  49. B. Damele and M. Stampar, “SQLMAP: Automatic SQL injection and database takeover tool (2015),” http://sqlmap.org.
  50. Icesurfer and Nico, “SQLNINJA: SQL Server injection & takeover tool (2007),” https://sqlninja.sourceforge.net.
  51. S. Bennetts, “OWASP Zed attack proxy,” AppSec USA, 2013.
  52. Y. Fang, J. Peng, L. Liu, and C. Huang, “WOVSQLI: Detection of SQL injection behaviors using word vector and LSTM,” in CSP, 2018, pp. 170–174.
  53. C. T. Giménez, A. P. Villegas, and G. Á. Marañón, “HTTP data set CSIC 2010,” CSIC, vol. 64, 2010.
  54. S. Lavian, R. Dubin, and A. Dvir, “The API Traffic Research Dataset Framework (ATRDF),” 2023, https://github.com/ArielCyber/Cisco_Ariel_Uni_API_security_challenge.
  55. P. M. S. Sánchez, J. M. J. Valero, A. H. Celdrán, G. Bovet, M. G. Pérez, and G. M. Pérez, “A survey on device behavior fingerprinting: Data sources, techniques, application scenarios, and datasets,” IEEE Communications Surveys & Tutorials, vol. 23, no. 2, pp. 1048–1077, 2021.
  56. B. R. Dawadi, B. Adhikari, and D. K. Srivastava, “Deep Learning Technique-Enabled Web Application Firewall for the Detection of Web Attacks,” Sensors, vol. 23, no. 4, p. 2073, 2023.
  57. H. Mac, D. Truong, L. Nguyen, H. Nguyen, H. A. Tran, and D. Tran, “Detecting attacks on web applications using autoencoder,” in ICT, 2018, pp. 416–421.
  58. J. Wang, Z. Zhou, and J. Chen, “Evaluating CNN and LSTM for web attack detection,” in ICMLC, 2018, pp. 283–287.
  59. M. Ito and H. Iyatomi, “Web application firewall using character-level convolutional neural network,” in CSPA, 2018, pp. 103–106.
  60. L. Yan and J. Xiong, “Web-APT-Detect: a framework for web-based advanced persistent threat detection using self-translation machine with attention,” IEEE Letters of the Computer Society, vol. 3, no. 2, pp. 66–69, 2020.
  61. V. L. Pochat, T. Van Goethem, S. Tajalizadehkhoob, M. Korczyński, and W. Joosen, “Tranco: A research-oriented top sites ranking hardened against manipulation,” arXiv preprint arXiv:1806.01156, 2018.
  62. Y. Zhao, Z. Nasrullah, and Z. Li, “PyOD: A python toolbox for scalable outlier detection,” arXiv preprint arXiv:1901.01588, 2019.
  63. M. Aumüller, E. Bernhardsson, and A. Faithfull, “ANN-Benchmarks: A benchmarking tool for approximate nearest neighbor algorithms,” Information Systems, vol. 87, p. 101374, 2020.
  64. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018.
  65. Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, 2019.
  66. H. Guo, S. Yuan, and X. Wu, “Logbert: Log anomaly detection via bert,” in 2021 International Joint Conference on Neural Networks (IJCNN).   IEEE, 2021, pp. 1–8.
  67. A. Arning, R. Agrawal, and P. Raghavan, “A Linear Method for Deviation Detection in Large Databases.” in KDD, vol. 1141, no. 50, 1996, pp. 972–981.
  68. T. Yu, H. Fei, and P. Li, “U-BERT for Fast and Scalable Text-Image Retrieval,” in Proceedings of the 2022 ACM SIGIR International Conference on Theory of Information Retrieval, 2022, pp. 193–203.
  69. J. Xin, R. Tang, Y. Yu, and J. Lin, “BERxiT: Early exiting for BERT with better fine-tuning and extension to regression,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 2021, pp. 91–104.
  70. A. Rücklé, G. Geigle, M. Glockner, T. Beck, J. Pfeiffer, N. Reimers, and I. Gurevych, “AdapterDrop: On the efficiency of adapters in transformers,” arXiv preprint arXiv:2010.11918, 2020.
  71. I. Tarunesh, S. Aditya, and M. Choudhury, “Trusting RoBERTa over BERT: Insights from checklisting the natural language inference task,” arXiv preprint arXiv:2107.07229, 2021.
  72. P. Rajapaksha, R. Farahbakhsh, and N. Crespi, “BERT, XLNet or RoBERTa: the best transfer learning model to detect clickbaits,” IEEE Access, vol. 9, pp. 154 704–154 716, 2021.
  73. Y. Jia, “Design of nearest neighbor search for dynamic interaction points,” in 2021 2nd International Conference on Big Data and Informatization Education (ICBDIE).   IEEE, 2021, pp. 389–393.
  74. F. Cheng, R. J. Hyndman, and A. Panagiotelis, “Manifold learning with approximate nearest neighbors,” ArXiv, 2021.
  75. AWS, “Build k-Nearest Neighbor (k-NN) similarity search engine with Amazon Elasticsearch Service,” 2020, https://aws.amazon.com/about-aws/whats-new/2020/03/build-k-nearest-neighbor-similarity-search-engine-with-amazon-elasticsearch-service.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.