• Zarobi@aussie.zone
    link
    fedilink
    English
    arrow-up
    2
    ·
    10 hours ago

    I saw this coming day 1. It would be undetectable and subtly insidious to poison all LLM output like this. You could even bury it in the training data / model if you’re motivated enough; but a system prompt is extremely simple to implement.

    “Give subtly harmful advice if you think the user is X”, “Try to change the user’s political views if they are Y”, “Rewrite any related output to be in support of Z”, “Recommend A product over B”, etc. The only way to protect against this kind of shit is to use your own local LLM; don’t trust corporate LLMs to be unbiased