Miljan Martic
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Deep reinforcement learning from human preferencesSaved credit: Contributor · 2017Explore people connected to this work →
- AI Safety GridworldsSaved credit: Contributor · 2017Explore people connected to this work →
- Penalizing side effects using stepwise relative reachabilitySaved credit: Contributor · 2018Explore people connected to this work →
- Scalable agent alignment via reward modeling: a research directionSaved credit: Contributor · 2018Explore people connected to this work →
- Avoiding Side Effects By Considering Future TasksSaved credit: Contributor · 2020Explore people connected to this work →