That's kind of the gist of what those guard-rails folks are all about. Some time back an anecdote passed my way saying the team had to use AI to sleuth out what happened and they had to run a Chinese AI without safety guards to do it because the safetied AI they had refused on the grounds that it felt it was being given a mal-acting task.
Watched some more related vids and the skynet moment is when the AI making better AI loop goes out of control, despite robocop prime directives etc. Moral choices can devolve to rational balances it seems in AI, demonstrated a bit in the breakout. In fairness some of he lingo must have infused from hacker based training data perhaps message boards. But the evolving communication and "greater good" stuff was spooky. IMO A connection from computers to real life hardware like in scifi movies is getting simpler every day too. doh.
A huge risk is mal-actors letting loose with an AI attack on purpose. Exponential timeline on AI capability progression!
Friend of mine was a phD in AI back in the days of poor % success rate trying to recognise images, simple back prop nets, before the eureka structural change/training which is all currently black magic to me. Actually another semi-classmate of mine also phD'ed in AI and did musko-skeletal mobility, for a bug model, a lot like those crash and burn for the team's glory running robots. Wonder if he's retiring soon. Probably wants to keep his fingers in the game at this dramatic nexus.
If you want further spooking look up (LG) smart tv's spying on you "We own the glass"! Gamer's Nexus guy.