Finite Sample Analysis of Average-Reward TD Learning and $Q$-Learning

Sheng Zhang, Zhe Zhang, Siva Theja Maguluri. Finite Sample Analysis of Average-Reward TD Learning and $Q$-Learning. In Marc'Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, Jennifer Wortman Vaughan, editors, Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual. pages 1230-1242, 2021. [doi]

@inproceedings{ZhangZM21-8,
  title = {Finite Sample Analysis of Average-Reward TD Learning and $Q$-Learning},
  author = {Sheng Zhang and Zhe Zhang and Siva Theja Maguluri},
  year = {2021},
  url = {https://proceedings.neurips.cc/paper/2021/hash/096ffc299200f51751b08da6d865ae95-Abstract.html},
  researchr = {https://researchr.org/publication/ZhangZM21-8},
  cites = {0},
  citedby = {0},
  pages = {1230-1242},
  booktitle = {Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual},
  editor = {Marc'Aurelio Ranzato and Alina Beygelzimer and Yann N. Dauphin and Percy Liang and Jennifer Wortman Vaughan},
}