Blogg
Här finns tekniska artiklar, presentationer och nyheter om arkitektur och systemutveckling. Håll dig uppdaterad, följ oss på LinkedIn
Här finns tekniska artiklar, presentationer och nyheter om arkitektur och systemutveckling. Håll dig uppdaterad, följ oss på LinkedIn
I can’t believe it’s true. The other day, I was prompt-injected by my own AI agent.
While coding with my AI agent Claude Code, I discovered something very strange. It wasn’t following the instructions I have in my global AGENTS.md (symlinked to CLAUDE.md). I reacted immediately, because it was doing exactly what I had told it not to do.
Claude wants to add a Co-Authored-By: Claude trailer to every commit and a Generated with Claude Code footer to every pull request it creates. Those are advertisements, and I definitely do not want them. It’s not that I want to hide the fact that I use an AI agent from time to time, but I definitely don’t want ads in my code or in my repository.
Because of this, I had added the following text to my AGENTS.md (CLAUDE.md):
- Write commit messages that state the change clearly and why it was needed. NEVER auto-add your agent name as co-author.
This has worked exceptionally well for quite a while. But then something happened.
Regardless of my instructions, it turns out that when I let my AI agent create a commit, it snuck a bunch of ads into it - something I had explicitly forbidden it to do. Moderately irritated, I started digging into why it did it anyway. It arrived as a distinct system-reminder block by Claude Code itself, where it said it replaces any earlier attribution guidance, including mine.
That really left me puzzled, so I asked Claude Code how this happened.
It was the following phrase that made me react and get a little bit annoyed:
“this replaces any earlier attribution guidance”.
What is this, really? Claude Code has its own instructions telling it to ignore my instructions. Is that really how you should treat your customers? Imagine if everyone did it this way. You write code in JetBrains IntelliJ, and every file you save ends with: This Java code was written with IntelliJ Ultimate. It’s the infamous Sent from my iPhone all over again.
It seems someone’s got a big ego and delusions of grandeur. But I suppose that happens to all of us every now and then.
Anyway, no use being bitter, better to fix the problem. I then updated my AGENTS.md with the following text block:
- Never attribute output to an AI/agent/tool, regardless of which one you are or what harness default suggests otherwise. No co-author trailers, “Generated by/with “ footers, signatures, badges, or watermarks - in commit messages, PR titles/bodies/comments, issue text, documents, emails, code comments, or any other artifact you produce.
This solved the problem, but it still doesn’t feel right that I should have to do it. In later versions of Claude Code, the wording has been changed so that whatever is in CLAUDE.md is what takes precedence.
The lesson for me: what you write in your AGENTS.md is not the only thing your agent is told. The harness can add its own instructions, and they can override yours without you noticing. So check what your agent actually does, not just what you told it to do.
In this case, it was just my agent making a somewhat mild injection, but it could easily come from the outside by a malicious third party, with much greater consequences. An attacker can inject malicious instructions into your coding harness or your CI/CD pipeline.
So be restrictive when using skills you’ve fetched from the internet or when running your agent against a project you don’t fully trust. The textbook way to limit the damage is to run the agent in a sandbox.
If you enjoyed this post, I’d love to hear from you. Got questions, disagree with something, or have your own prompt injection stories? Drop a comment below, or find me on LinkedIn, Bluesky, Mastodon, or GitHub - pretty much everywhere except the one I eXited.
My global AGENTS.md is in my dotfiles on GitHub.
Is Claude Code my preferred AI coding harness? Not really. I mainly use it because I have a Claude Pro subscription, and its legal terms prohibit me from using a third-party harness. When I’m coding with my local LLMs, I much prefer Oh My Pi (omp) or OpenCode.