Papers
Topics
Authors
Recent
Search
2000 character limit reached

Navigation as Attackers Wish? Towards Building Robust Embodied Agents under Federated Learning

Published 27 Nov 2022 in cs.AI, cs.CL, cs.CR, and cs.CV | (2211.14769v4)

Abstract: Federated embodied agent learning protects the data privacy of individual visual environments by keeping data locally at each client (the individual environment) during training. However, since the local data is inaccessible to the server under federated learning, attackers may easily poison the training data of the local client to build a backdoor in the agent without notice. Deploying such an agent raises the risk of potential harm to humans, as the attackers may easily navigate and control the agent as they wish via the backdoor. Towards Byzantine-robust federated embodied agent learning, in this paper, we study the attack and defense for the task of vision-and-language navigation (VLN), where the agent is required to follow natural language instructions to navigate indoor environments. First, we introduce a simple but effective attack strategy, Navigation as Wish (NAW), in which the malicious client manipulates local trajectory data to implant a backdoor into the global model. Results on two VLN datasets (R2R and RxR) show that NAW can easily navigate the deployed VLN agent regardless of the language instruction, without affecting its performance on normal test sets. Then, we propose a new Prompt-Based Aggregation (PBA) to defend against the NAW attack in federated VLN, which provides the server with a ''prompt'' of the vision-and-language alignment variance between the benign and malicious clients so that they can be distinguished during training. We validate the effectiveness of the PBA method on protecting the global model from the NAW attack, which outperforms other state-of-the-art defense methods by a large margin in the defense metrics on R2R and RxR.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (37)
  1. On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757.
  2. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  3. How to backdoor federated learning. In International Conference on Artificial Intelligence and Statistics, pages 2938–2948. PMLR.
  4. signsgd with majority vote is communication efficient and fault tolerant. arXiv preprint arXiv:1810.05291.
  5. Analyzing federated learning through an adversarial lens. In International Conference on Machine Learning, pages 634–643. PMLR.
  6. Machine learning with adversaries: Byzantine tolerant gradient descent. In NIPS.
  7. Fltrust: Byzantine-robust federated learning via trust bootstrapping. ArXiv, abs/2012.13995.
  8. Robustnav: Towards benchmarking robustness in embodied navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15691–15700.
  9. Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint arXiv:1712.05526.
  10. Badnets: Identifying vulnerabilities in the machine learning model supply chain. ArXiv, abs/1708.06733.
  11. Promptfl: Let federated participants cooperatively learn prompts instead of models - federated learning in age of foundation model. ArXiv, abs/2208.11625.
  12. Towards learning a generic agent for vision-and-language navigation via pre-training. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  13. Cpl: Counterfactual prompt learning for vision and language models. arXiv preprint arXiv:2210.10362.
  14. Vln bert: A recurrent vision-and-language bert for navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1643–1653.
  15. Adversarial machine learning. In Proceedings of the 4th ACM workshop on Security and artificial intelligence, pages 43–58.
  16. Communication-efficient distributed sgd with sketching. Advances in Neural Information Processing Systems, 32.
  17. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4392–4412.
  18. Stacked cross attention for image-text matching. In Proceedings of the European conference on computer vision (ECCV), pages 201–216.
  19. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXX, volume 12375 of Lecture Notes in Computer Science, pages 121–137. Springer.
  20. Spatiotemporal attacks for embodied agents. In European Conference on Computer Vision, pages 122–138. Springer.
  21. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  22. Gpt understands, too. arXiv:2103.10385.
  23. Privacy and robustness in federated learning: Attacks and defenses. ArXiv, abs/2012.06337.
  24. The hidden vulnerability of distributed learning in byzantium. In ICML.
  25. Asynchronous methods for deep reinforcement learning. In International conference on machine learning, pages 1928–1937. PMLR.
  26. REVERIE: remote embodied visual referring expression in real indoor environments. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 9979–9988. Computer Vision Foundation / IEEE.
  27. Learning transferable visual models from natural language supervision. In ICML.
  28. Timo Schick and Hinrich Schütze. 2020. Exploiting cloze questions for few shot text classification and natural language inference. arXiv preprint arXiv:2001.07676.
  29. How much can CLIP benefit vision-and-language tasks? In International Conference on Learning Representations.
  30. Learning to navigate unseen environments: Back translation with environmental dropout. arXiv preprint arXiv:1904.04195.
  31. Decentralized collaborative learning of personalized models over networks. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, AISTATS 2017, 20-22 April 2017, Fort Lauderdale, FL, USA, Proceedings of Machine Learning Research, pages 509–517.
  32. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6629–6638.
  33. Dba: Distributed backdoor attacks against federated learning. In International Conference on Learning Representations.
  34. Byzantine-robust distributed learning: Towards optimal statistical rates. ArXiv, abs/1803.01498.
  35. Neurotoxin: Durable backdoors in federated learning. In International Conference on Machine Learning, pages 26429–26446. PMLR.
  36. Kaiwen Zhou and Xin Eric Wang. 2022. Fedvln: Privacy-preserving federated vision-and-language navigation. arXiv preprint arXiv:2203.14936.
  37. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16816–16825.
Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.