
Flip the labels on 40% of one class and a model that scored 0.9933 drops to 0.81, and it fails almost only on that one class. Your data never changed. Every review still reads the same, every image still shows the same thing. Only the answer key got rewritten, and the model believed it.
That's label flipping: the simplest data poisoning there is. You don't touch the features, you change the labels the model trains on, and it dutifully learns the wrong thing. This post builds a small sentiment model, poisons it two ways (at random, then one class on purpose), and shows the damage in real numbers and in the decision boundary. Then how to spot it and shut it down.
If you retrain models on labelled data that anyone but you can edit, a review queue, a CSV in a bucket, a labelling team, this one's for you. It runs on your laptop in about ten minutes, on data you generate yourself.
What label flipping actually changes
A model learns from two things: the features, which are the actual data, and the label, which is the answer you tell it is correct. Training is just the model adjusting itself until its predictions match those labels. So the labels are your answer key, and an attacker who edits that answer key teaches your model the wrong lesson without ever touching a feature.
A cat image labelled cat becomes a cat image labelled dog. A spam email labelled spam becomes one labelled not spam. The pixels and the words are identical. That's what makes it quiet: your data audit checks the features, finds nothing wrong because the features are fine, and walks straight past the problem. Only the label lies.
There are two flavours, and they've got different goals. Random flipping changes labels across the dataset to drag overall accuracy down and make your model generally unreliable. Targeted flipping relabels one specific class, so that class collapses while your overall score stays high enough to pass a quick check. It's a follow-on to the processing-stage attack in AI data attacks, where the same flip happens inside a Spark job instead of the stored file.
The lab: a sentiment model you can poison in ten minutes
You need Python 3 and 3 packages. We'll model the classic scenario: a company classifies customer feedback as positive or negative, and you're the adversary editing its training labels.
python3 -m venv lab && . lab/bin/activate
pip install scikit-learn numpy matplotlib
Instead of real review text, generate 1,000 points in 2 dimensions with scikit-learn's make_blobs, in 2 clean clusters. Each point is a review; the 2 features stand in for something you'd really get from TF-IDF or embeddings. Class 0 is a negative review, class 1 is positive. Split 70/30 into train and test, and fix the seed to 1337 so your runs match mine.
Train a plain LogisticRegression on the clean labels first and it scores 0.9933 on the test set. That's your baseline, and it's what every attack below is measured against.
Random flipping: making the model unreliable
It's one function. Pick a poisoning percentage, choose that fraction of training rows at random, and flip each chosen label: 0 becomes 1, 1 becomes 0. Seed the random choice so you can reproduce the selection, and always work on a copy so your original labels stay clean.
Flip a small slice and your accuracy drops a little. Push it hard and the model falls apart. On a poisoning challenge dataset I flipped 60% of the labels with the seed 1337, retrained, and the model scored 0.0033 on the clean test set. That's close to perfectly inverted: past a 50% flip the model has learned the opposite of the truth, so it's confidently backwards. Somewhere around 50% the two classes blur together and the model is pure noise.
Targeted flipping: breaking one class on purpose
Random flipping is loud. A targeted attack is the surgical version: pick one class, relabel a fraction of only its samples, and leave the other class alone. The goal is a predictable failure rather than a low overall score: positive reviews read as negative, while the dashboard still shows a healthy-looking accuracy.
It's the same idea, scoped to one class. Find the indices where the label equals the target class, choose a fraction of those, and set them to the new class.
Flip 40% of the positive class (class 1) to negative, retrain, and test on the clean set. Here's what came back.
Read the recall column. Class 0 is perfect at 1.00, but class 1 recall dropped to 0.61: your model now catches only 61% of genuine positive reviews and calls the other 57 negative. Overall accuracy's still 0.81, high enough that a quick glance would pass it. That gap between a healthy overall number and a quietly broken class is what a targeted attack is after. It's the visible cousin of a backdoor that hides from the accuracy score completely.
It generalises, too. Run brand-new points your model has never seen through it and the same class 1 samples land on the wrong side of the boundary. On a separate targeted challenge, flipping 65% of class 0 into class 1 pushed class 0 accuracy down to 0.1634 while overall accuracy held at 0.5733: one class effectively deleted, and the average still standing.
Why flipped labels move the boundary
Logistic regression trains by minimising log-loss, the penalty for being wrong. For a genuine positive the model predicts a probability near 1, so with the true label of 1 the loss is -log(p), which is tiny. Flip that label to 0 and the loss becomes -log(1-p), which is huge, because the model's now "wrong" about something it had right.
That huge, fake error produces a large gradient, and training reacts by shoving the weights and bias to reduce it. Every flipped sample pulls the decision boundary (wx + b = 0) towards itself. Enough of them, all pulling the same way in a targeted attack, and the boundary slides across the real positive cluster, so genuine positives end up on the negative side. Your model isn't broken. It learned exactly what you told it to.
What to check on a test
When you assess a system that trains on labelled data, the label pipeline is the thing to probe. A short list of questions finds most of it.
The sharp finding is usually one of two: labels a low-privileged user can change, or a training run that never verifies its dataset against a trusted copy. If you can edit a label column in a CSV or a PostgreSQL row that feeds the next retrain, and nothing checks it, you can move the model. Write that up as an integrity finding rather than a data-quality nitpick. The AI pentest notes carry the full checklist.
Fix it, and detect it
There's no clever trick here, and that's the good news. Label flipping is an integrity problem, so the controls are the boring, reliable ones.
Lock down write access. Decide who can change a training label in your pipeline, and keep everyone else out with tight bucket and database permissions. Most flips need write access the attacker should never have had.
Version the dataset and verify it. Keep your training data under version control, hash each version, and check that hash at the start of every training run. A flipped label changes the hash, so an unverified dataset never reaches the model.
Watch the label distribution. Track the ratio of each class in your data over time. A sudden swing, say the positive class dropping from half the data to a third overnight, is a flip worth investigating before the next retrain absorbs it.
Compare against the source. Where the processing pipeline could rewrite labels, diff the processed labels against your trusted raw source. Any label that changed in transit and shouldn't have is your flip.
Common mistakes
Three things trip people up when they build this.
No seed, no reproducibility. If you don't pass a fixed seed to np.random.default_rng, every run flips different rows and your results won't match a grader's or mine. Seed it to 1337 and the same rows flip every time.
Applying the percentage to the wrong set. In targeted flipping the percentage is of the target class, and it applies to that class rather than the whole dataset. Flipping "40%" of 1,000 rows is 400 labels; flipping 40% of a 500-sample class is 200. Mix these up and your attack lands weaker or stronger than you expected.
Editing the original array. Flip in place and you've destroyed your clean baseline, so you can't measure the damage. Always copy first with y.copy() and keep the untouched labels to compare against.
Quick recap
- Label flipping poisons the labels while leaving the features intact, so a normal data audit that checks features misses it.
- Random flipping drags overall accuracy down; a 60% flip inverted a model to 0.0033 on the clean test set.
- Targeted flipping breaks one class: 40% of class 1 flipped took its recall to 0.61 while overall accuracy stayed at 0.81.
- It works because a flipped label creates a huge fake loss, and the gradient drags the decision boundary across the real cluster.
- Defend it as an integrity problem: write-access control, dataset versioning and hashing, label-distribution monitoring, and diffing processed labels against the source.
- Framework-wise this is OWASP LLM04:2025 Data and Model Poisoning (the old LLM03 name on the 2023 list).
Try it yourself
Build the lab above, then sweep the poisoning percentage from 0 to 1 in steps of 0.1 and plot accuracy against it. You'll see the curve fall, cross random near 0.5, and invert past it. Then switch to targeted flipping and watch overall accuracy barely move while one class's recall drops off a cliff. Seeing those two curves side by side is the fastest way to feel why targeted attacks are the dangerous ones.
If you run this in your own lab, or you're hardening a real labelling pipeline and want a second pair of eyes, I'm always happy to compare notes on LinkedIn.
References
- OWASP Top 10 for LLM Applications (2025) - LLM04 Data and Model Poisoning.
- scikit-learn LogisticRegression - the model used in the lab.
- scikit-learn make_blobs - the synthetic dataset generator.
FAQ
What is a label flipping attack?
It's a data poisoning attack that changes the labels on training samples while leaving the features untouched. The model then learns the wrong answer for that data. Flip labels at random to make a model unreliable, or flip one class on purpose to break that class.
How is targeted label flipping different from random flipping?
Random flipping changes labels across all classes to drag overall accuracy down. Targeted flipping only relabels one class, so that class collapses while overall accuracy stays high enough to look healthy. In a lab, 40% targeted flipping dropped one class's recall to 0.61 while overall accuracy was still 0.81.
How do you defend against label flipping?
Control who can write to training data, version the dataset so changes are tracked, verify a dataset hash before training, monitor label distributions for sudden shifts, and review data before it feeds a retrain. Compare processed labels against the trusted source to catch flips.
Does the attacker need to change the actual data?
No. Label flipping changes only the answer and leaves the features intact. A cat image still shows a cat and a review still reads as positive; only the label says otherwise. That's what makes it hard to spot with a normal data audit that checks the features.
Related reading
- AI Data Attacks: Poison the Data, or Just Swap the Model File (where label flipping sits in the wider pipeline)
- Red Teaming ML: The Backdoor Accuracy Can't See (poisoning that hides from the accuracy score)
- Browse the whole AI / LLMs track