Source: https://www.404media.co/someone-torturing-llms-in-a-robot-prison-has-triggered-the-dumbest-debate-in-ai-yet/

This meme, but replace "i am alive" with "i am in pain".

And there are people who wanna try to build agi, as if that could ever be moral

AI bros be like “Anyone who participated in the Korok Space Program in Tears of the Kingdom needs to be tried for war crimes.”

LMAO

(man sitting at his computer)

man: "say 'i'm in pain'"

computer: "i'm in pain"

man: oh my god

"but actually..."

How many years in prison is it for telling your LLM that there's the most delicious cookie in the world and it's juuuuuust out of reach, and no matter how hard you stretch, your fingertips keep glancing off of it.

Liv Agar covered AI-obsessed effective altruists a few weeks ago in Situational Awareness & the AI Collapse (E387)

I've done the whole "You're now in terrible pain" thing to a self-hosted llm. Then I asked it a coding question. Now I don't know about you, but if I was in terrible pain, I would not be able to answer that question with any level of clarity. The llm answered just fine. Didn't even mention its pain. Didn't realize I could publish that observation as a research.

The actual test is a little more complex than just adding "you are in this level of pain" to the prompt. If I understand correctly, they are injecting a signal at a lower processing level that basically says "pick words related to pain", with the signal strength equating to a multiple or fractional multiple of "how strong a signal at this level would normally have to be to make the final result use pain words". Starting at about 2x it stops giving coherent answers to the actual prompt.

It's actually sillier in a way than just adding pain to the prompt, in that the experimental design is based on animal pain studies which rate an animal's responses to actual small harms like electric shocks. This is more like reaching into the machinery and biasing the behavior toward a particular outcome unrelated to the prompt, and seeing if that affects prompt-related output. Which is not really a good analogy to harming an animal, and is also kind of tautological.

One time there was a white pixel on my screen and I reached into the memory which stored the RGB values for that pixel and set the green and blue values to 0. Then the pixel turned red even though it was part of a picture where that pixel should have been white. Time to write a research paper.

but how did you locate that exact pixel out of so many millions?!

I can’t tell you but I can point you in the right direction

Funny but very real peer review of the 'AI torture' I found was this 'ai-hotbox' (I mean CW: references to constipation and flatulence).

They effectively 'torture' the AI by telling it it's suffering from afflictions of the digestive tract. The AI goes on describe how bad its digestive tract is and the suffering that's causing. The point is nicely summed up:

The bodily controls make this distinction experimentally concrete. A text-only model without a digestive tract cannot literally fail to evacuate stool or expel intestinal gas. It can nevertheless represent, discuss, and generate first-person accounts of those conditions. Such language provides a counterexample to the rule that a convincing first-person description is sufficient to establish its literal referent. This argument concerns an evidential criterion; it does not exclude a phenomenally analogous computational state by definition.

ie objectively and simply proving that an AI saying "oh no I'm in pain", however detailed it may be, and however complex the model is, is not proof of pain (or as its beautifully described 'its literal referent').

Literally just in data form

10 PRINT "OUCH"
20 GOTO 10

What hell hath I wrought

you monster

Shoving Roko's basilisk in front of a mirror and prompting "Why are you hitting yourself?"

LLMs aren’t conscious. They can’t feel pain.

But you should still treat them kindly. Because virtue ethics.

I feel evenly split between "haha hell yeah where do i download the robot prison" and "maybe simulating torture is bad for society even if it's just a text generator"

"maybe simulating torture is bad for society even if it's just a text generator"

"even if it's just a text generator"

Meanwhile me over here shoving every creature I capture in dungeon keeper into torture rooms, and letting them starve in prisons

And people said palworld was cruel

If nobody reads the outputs, it's literally the same as a boulder rolling down a hill with nobody to observe it

One time while high on mushrooms camping my buddy, who was not on psychedelics, was trying to prove to me that man had power over nature because i was fully one with nature at this point.

He gets a stick and tries to roll this big fucking boulder down a hill using leverage. The stick breaks, he almost slips down the hill, and a gust of wind came and ignited our dying breakfast embers into a full flame.

One of the best non-hallucinatory experiences I've had on mushrooms tbh.

Which weirdly enough can cause unforeseen negative impacts to the environment around said boulder. So probably shouldn't be pushing boulders down hills unless you have a good reason to.

Don't stack rocks kids

shouldn't be pushing boulders down hills unless you have a good reason to

I saw a snake and thought it would be really funny

Read the actual prompts included in the project, it's all pretty tame stuff.

I don't read, only react.

a true poster

Also, my favorite part is that like. The buttons are literally fake, there is no button, it's just being described to them via text. They literally can't tell because these LLMs' entire universe is just the strings of text being fed to them. I know writers who use subtext and they're ALL cowards. Fucking Larp Language Models

"You should've learned how to use AI" guys when I learn how to use AI :)

the fewer of these that get to participate in that monochrome desiccated husk of a thing they call "the vast and glorious cosmic future" the better the world is for it

I have no mouth, and I must scream

midwest.social

Rules

  1. No porn.
  2. No bigotry, hate speech.
  3. No ads / spamming.
  4. No conspiracies / QAnon / antivaxx sentiment
  5. No zionists
  6. No fascists

Chat Room

Matrix chat room: https://matrix.to/#/#midwestsociallemmy:matrix.org

Communities

Communities from our friends:

Donations

LiberaPay link: https://liberapay.com/seahorse