Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement Learning
Abstract: Offline safe reinforcement learning (RL) aims to train a constraint satisfaction policy from a fixed dataset. Current state-of-the-art approaches are based on supervised learning with a conditioned policy. However, these approaches fall short in real-world applications that involve complex tasks with rich temporal and logical structures. In this paper, we propose temporal logic Specification-conditioned Decision Transformer (SDT), a novel framework that harnesses the expressive power of signal temporal logic (STL) to specify complex temporal rules that an agent should follow and the sequential modeling capability of Decision Transformer (DT). Empirical evaluations on the DSRL benchmarks demonstrate the better capacity of SDT in learning safe and high-reward policies compared with existing approaches. In addition, SDT shows good alignment with respect to different desired degrees of satisfaction of the STL specification that it is conditioned on.
- Constrained policy optimization. In International conference on machine learning, pp. 22–31. PMLR, 2017.
- Altman, E. Constrained Markov decision processes. Routledge, 2021.
- Concrete problems in ai safety. arXiv preprint arXiv:1606.06565, 2016.
- Structured reward shaping using signal temporal logic specifications. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 3481–3486. IEEE, 2019.
- Control synthesis from linear temporal logic specifications using model-free reinforcement learning. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pp. 10349–10355. IEEE, 2020.
- When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35:1542–1553, 2022.
- Safe learning in robotics: From learning-based control to safe reinforcement learning. Annual Review of Control, Robotics, and Autonomous Systems, 5:411–444, 2022.
- Ltl and beyond: Formal languages for reward function specification in reinforcement learning. In IJCAI, volume 19, pp. 6065–6073, 2019.
- A near-optimal primal-dual method for off-policy learning in cmdp. Advances in Neural Information Processing Systems, 35:10521–10532, 2022.
- Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34:15084–15097, 2021.
- Lyapunov-based safe policy optimization for continuous control. arXiv preprint arXiv:1901.10031, 2019.
- Conditional positional encodings for vision transformers. arXiv preprint arXiv:2102.10882, 2021.
- Model-based reinforcement learning for approximate optimal control with temporal logic specifications. In Proceedings of the 24th International Conference on Hybrid Systems: Computation and Control, pp. 1–11, 2021.
- Formal methods for control of traffic flow: Automated control synthesis from finite-state transition models. IEEE Control Systems Magazine, 37(2):109–128, 2017.
- Robust counterexample-guided optimization for planning from differentiable temporal logic. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 7205–7212. IEEE, 2022.
- Robust online monitoring of signal temporal logic. Formal Methods in System Design, 51:5–30, 2017.
- On-line monitoring for temporal logic robustness. In International Conference on Runtime Verification, pp. 231–246. Springer, 2014.
- Donzé, A. On signal temporal logic. In Runtime Verification: 4th International Conference, RV 2013, Rennes, France, September 24-27, 2013. Proceedings 4, pp. 382–383. Springer, 2013.
- Robust satisfaction of temporal logic over real-valued signals. In International Conference on Formal Modeling and Analysis of Timed Systems, pp. 92–106. Springer, 2010.
- Efficient robust monitoring for stl. In Computer Aided Verification: 25th International Conference, CAV 2013, Saint Petersburg, Russia, July 13-19, 2013. Proceedings 25, pp. 264–279. Springer, 2013.
- Objective functions for falsification of signal temporal logic properties in cyber-physical systems. In 2017 13th IEEE Conference on Automation Science and Engineering (CASE), pp. 1326–1331. IEEE, 2017.
- Rvs: What is essential for offline rl via supervised learning? arXiv preprint arXiv:2112.10751, 2021.
- Robustness of temporal logic specifications for continuous-time signals. Theoretical Computer Science, 410(42):4262–4291, 2009.
- D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020.
- Off-policy deep reinforcement learning without exploration. corr abs/1812.02900 (2018). arXiv preprint arXiv:1812.02900, 2018.
- Generalized decision transformer for offline hindsight information matching. In International Conference on Learning Representations, 2022.
- Gronauer, S. Bullet-safety-gym: Aframework for constrained reinforcement learning. 2022.
- Constrained reinforcement learning for vehicle motion planning with topological reachability analysis. Robotics, 11(4):81, 2022a.
- A review of safe reinforcement learning: Methods, theory and applications. arXiv preprint arXiv:2205.10330, 2022b.
- Kat: A knowledge augmented transformer for vision-and-language. arXiv preprint arXiv:2112.08614, 2021.
- Sample efficient offline-to-online reinforcement learning. IEEE Transactions on Knowledge and Data Engineering, 2023.
- A primal-dual-critic algorithm for offline constrained reinforcement learning. arXiv preprint arXiv:2306.07818, 2023.
- Graph decision transformer. arXiv preprint arXiv:2303.03747, 2023.
- Omnisafe: An infrastructure for accelerating safe reinforcement learning research. arXiv preprint arXiv:2305.09304, 2023.
- Model-based reinforcement learning from signal temporal logic specifications. arXiv preprint arXiv:2011.04950, 2020.
- Towards safe mechanical ventilation treatment using deep offline reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp. 15696–15702, 2023.
- Offline reinforcement learning with fisher divergence critic regularization. In International Conference on Machine Learning, pp. 5774–5783. PMLR, 2021.
- Stabilizing off-policy q-learning via bootstrapping error reduction. Advances in Neural Information Processing Systems, 32, 2019a.
- Reward-conditioned policies. arXiv preprint arXiv:1912.13465, 2019b.
- A more scalable mixed-integer encoding for metric temporal logic. IEEE Control Systems Letters, 6:1718–1723, 2021.
- Batch policy learning under constraints. In International Conference on Machine Learning, pp. 3703–3712. PMLR, 2019.
- Optidice: Offline policy optimization via stationary distribution correction estimation. In International Conference on Machine Learning, pp. 6120–6130. PMLR, 2021.
- Coptidice: Offline constrained reinforcement learning via stationary distribution correction estimation. arXiv preprint arXiv:2204.08957, 2022.
- Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
- Guided online distillation: Promoting safe reinforcement learning by offline demonstration. arXiv preprint arXiv:2309.09408, 2023.
- Reinforcement learning with temporal logic rewards. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 3834–3839. IEEE, 2017.
- Safety-aware causal representation for trustworthy reinforcement learning in autonomous driving. arXiv preprint arXiv:2311.10747, 2023.
- On the robustness of safe reinforcement learning under observational perturbations. arXiv preprint arXiv:2205.14691, 2022.
- Datasets and benchmarks for offline safe reinforcement learning. arXiv preprint arXiv:2306.09303, 2023a.
- Constrained decision transformer for offline safe reinforcement learning. arXiv preprint arXiv:2302.07351, 2023b.
- Metrics for signal temporal logic formulae. In 2018 IEEE Conference on Decision and Control (CDC), pp. 1542–1547. IEEE, 2018.
- Monitoring temporal properties of continuous signals. In International Symposium on Formal Techniques in Real-Time and Fault-Tolerant Systems, pp. 152–166. Springer, 2004.
- You can’t count on luck: Why decision transformers and rvs fail in stochastic environments. Advances in Neural Information Processing Systems, 35:38966–38979, 2022.
- Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019.
- Constrained offline policy optimization. In International Conference on Machine Learning, pp. 17801–17810. PMLR, 2022.
- A survey on offline reinforcement learning: Taxonomy, review, and open problems. IEEE Transactions on Neural Networks and Learning Systems, 2023.
- Improving language understanding by generative pre-training. 2018.
- Benchmarking safe exploration in deep reinforcement learning. arXiv preprint arXiv:1910.01708, 7(1):2, 2019.
- Autonomous vehicle decision-making and monitoring based on signal temporal logic and mixed-integer programming. In 2020 American Control Conference (ACC), pp. 454–459. IEEE, 2020.
- Schmidhuber, J. Reinforcement learning upside down: Don’t predict rewards–just map them to actions. arXiv preprint arXiv:1912.02875, 2019.
- Responsive safety in reinforcement learning by pid lagrangian methods. In International Conference on Machine Learning, pp. 9133–9143. PMLR, 2020.
- Robust linear temporal logic. arXiv preprint arXiv:1510.08970, 2015.
- Teaching multiple tasks to an rl agent using ltl. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, pp. 452–461, 2018.
- Minimum-violation scltl motion planning for mobility-on-demand. In 2017 IEEE International Conference on Robotics and Automation (ICRA), pp. 1481–1488. IEEE, 2017.
- Attention is all you need. Advances in neural information processing systems, 30, 2017.
- Critic-guided decision transformer for offline reinforcement learning. arXiv preprint arXiv:2312.13716, 2023.
- Constraints penalized q-learning for safe offline reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 8753–8760, 2022.
- Crpo: A new approach for safe reinforcement learning with convergence guarantee. In International Conference on Machine Learning, pp. 11480–11491. PMLR, 2021.
- Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl. In International Conference on Machine Learning, pp. 38989–39007. PMLR, 2023.
- Balancing therapeutic effect and safety in ventilator parameter recommendation: An offline reinforcement learning approach. Engineering Applications of Artificial Intelligence, 131:107784, 2024.
- Saformer: A conditional sequence modeling approach to offline safe reinforcement learning. arXiv preprint arXiv:2301.12203, 2023.
- Model-based safe reinforcement learning with time-varying state and control constraints: An application to intelligent vehicles. arXiv preprint arXiv:2112.11217, 2021.
- Modularized control synthesis for complex signal temporal logic specifications. arXiv preprint arXiv:2303.17086, 2023.
- Explicit sparse transformer: Concentrated attention through explicit selection. arXiv preprint arXiv:1912.11637, 2019.
- Online decision transformer. In international conference on machine learning, pp. 27042–27059. PMLR, 2022.
- Programmatic reward design by example. Proceedings of the AAAI Conference on Artificial Intelligence, 36(8):9233–9241, Jun. 2022. doi: 10.1609/aaai.v36i8.20910.
Paper Prompts
Sign up for free to create and run prompts on this paper.