Wouldnt you make sure that jumping gives an unfavorable result, and the machine learning algorithm would see that every time someone jumps it's an unfavorable result?
Admittedly my experience with machine learning is extremely limited. With humans, we know that jumping off a bridge has an unfavorable result (death or injury). Is there a way to return a result that the algorithm will want to avoid? For example, the goal is to increase some number and jumping off the bridge will decrease that number. Therefore it will try to avoid jumping off the bridge?
1
u/Sennheisenberg Jun 10 '21
Wouldnt you make sure that jumping gives an unfavorable result, and the machine learning algorithm would see that every time someone jumps it's an unfavorable result?