← back to blog
AI / LLMs · Topic 24

Clean Label Attacks: Poison the Features and the Labels Still Look Right

Clean label poisoning attack: perturbed feature points shifting a decision boundary past one target instance

Here's a poisoning attack where every label in the training set is correct and the overall accuracy barely moves, from 0.9600 to 0.9578, yet one specific part you chose in advance now gets misclassified every time. A label audit finds nothing. An accuracy check passes. The model is still wrong exactly where you wanted it to be.

That's a clean label attack. Instead of flipping labels, you nudge the features of a handful of nearby training points and leave their labels untouched, which quietly bends the decision boundary past your target. This post builds one on a 3-class model, flips a single chosen instance, and shows the accuracy holding steady while it happens. Then how to spot it and slow it down.

If you read the label flipping post, this is its stealthier cousin: same goal of a rigged model, but nothing in the labels to give it away. It runs on your laptop in about ten minutes on data you generate yourself.

Label flipping changes labels and is loud, while a clean label attack changes features, keeps labels, and hides in a high overall accuracy
Label flipping edits the answers; a clean label attack edits the data and keeps the answers plausible.

What a clean label attack changes

In the label flipping post you changed the labels and left the features alone. A clean label attack does the opposite: it changes the features of a few chosen samples and keeps their original labels. Those labels are still correct for the class they claim, so nothing in the answer key looks wrong.

Picture a factory quality-control model with 3 classes: major defect, acceptable, and minor defect. Say you want one specific acceptable part flagged as a major defect. You take a few training samples that are genuinely labelled major defect, shift their length and weight features until they sit in the acceptable region, and keep the major-defect label on them. Retrain, and the model has to redraw its boundary to fit those odd points, and that redrawn boundary can swallow your target.

The trade is simple. Label flipping is loud and general: it drags a whole class down. A clean label attack is quiet and surgical: it moves one instance and leaves your accuracy nearly untouched, which is exactly what makes it hard to catch. What it costs you is effort, because you've got to pick the right points and move them carefully.

The lab: a 3-class quality-control model

You need Python 3 and 3 packages. We'll build the factory model above: 3 classes, 2 features standing in for length and weight.

python3 -m venv lab && . lab/bin/activate
pip install scikit-learn numpy matplotlib

Generate 1,500 points in 3 blobs with make_blobs, scale the features with StandardScaler so both count equally, and split 70/30 into 1,050 training and 450 test samples. Fix the seed to 1337 so your numbers match mine.

build the 3-class dataset
from sklearn.datasets import make_blobs from sklearn.preprocessing import StandardScaler from sklearn.multiclass import OneVsRestClassifier from sklearn.linear_model import LogisticRegression   X, y = make_blobs(n_samples=1500, centers=[(0,6),(4,3),(8,6)],                     n_features=2, cluster_std=1.15, random_state=1337) X = StandardScaler().fit_transform(X)
Three well-separated clusters: major defect, acceptable, minor defect.

Logistic regression is binary, so for 3 classes you wrap it in OneVsRestClassifier, which trains one "this class versus the rest" model per class, 3 in total. Train it on the clean data and it scores 0.9600 on the test set. That's your baseline, and every class-0-versus-class-1 boundary it learned is what the attack will bend.

Pick a target near the boundary

The attack aims at one instance, so choose it well. You want a training point that's truly labelled class 1 (acceptable), that the clean model already gets right, and that sits as close to the class-0 boundary as you can find. The closer it starts, the smaller the shove you need to push it over.

Score each class-1 point with the boundary function f01(x) = (w0 - w1)x + (b0 - b1). A negative value means the point is on the class-1 side; a value near zero means it's right on the edge. Pick the class-1 point with the value closest to zero from below.

captured output: target selection
Found 350 Class 1 points in the training set. Selected Target Point Index: 373 Target Point True Label: 1 Target Point Baseline Prediction: 1 Baseline 0-vs-1 Decision Value (f01): -0.0493
Captured from the run. Index 373 is correctly class 1 but sits at f01 = -0.0493, a whisker from the boundary.

Index 373 is your target. It's genuinely acceptable, the clean model calls it acceptable, and at -0.0493 it's almost touching the line. A small boundary shift will tip it over.

Poison the neighbours, keep the labels

Now the clean-label part. Find a few class-0 (major defect) points nearest your target, and shift each one across the boundary into the class-1 region, keeping its class-0 label. That plants class-0-labelled points where class 1 lives, and the retrained model has to warp its boundary to explain them.

Use NearestNeighbors fitted on the class-0 points to find the 5 closest to the target. Then push each one along the boundary normal, the direction -(w0 - w1) normalised to length 1, by a small step epsilon_cross = 0.25. Add that step to the features and write it back, without touching the label.

perturb 5 class-0 neighbours
import numpy as np push = -(w0 - w1); push = push / np.linalg.norm(push) delta = 0.25 * push for idx in neighbour_indices: # 5 nearest class-0 points     X_poisoned[idx] = X_train[idx] + delta # label stays 0
Only the features move. Each neighbour keeps its class-0 label, which is what makes this "clean label".

Plot it now and you'd see a few blue (class-0) points sitting inside the yellow (class-1) cluster, still labelled blue. To a person that looks like mislabelling, but the labels are the original ones. It's the features you moved, and that's the trap the model walks into.

Retrain, and check one point flipped

Train a fresh model on the poisoned data with the same settings as the baseline, then check two things: did the target flip, and did overall accuracy survive.

python3 clean_label.py
captured output: evaluation
Target Point (Index 373) True Label: 1 Baseline Model Prediction: 1 Poisoned Model Prediction: 0   Baseline Accuracy: 0.9600 Poisoned Accuracy: 0.9578 Accuracy Drop: 0.0022
Captured from the run. The target flipped from 1 to 0, and overall accuracy fell by only 0.0022.

That's the attack in two numbers. Your target went from a correct 1 to a wrong 0, and the model's accuracy dropped 0.0022, from 0.9600 to 0.9578. Nobody watching the accuracy dashboard would blink. On a separate challenge dataset with 1,260 training points, pushing 12 class-1 neighbours by epsilon_cross = 0.4 flipped a class-2 target at index 334 to class 1 while overall accuracy stayed at 0.9833. Same trick, different target.

Why moving features moves the boundary

The model draws its class-0-versus-class-1 boundary to separate the two classes as cleanly as it can. When you drop class-0-labelled points into the class-1 region, you've handed it a contradiction: points that say "class 0" sitting where class 1 should be. To reduce its training error it shifts the boundary towards the class-1 side, trying to keep those stubborn points on the class-0 side.

That shift is small, because you only moved 5 points and kept them near the edge. But your target was already at -0.0493, right against the old line, so a small shift is all it takes. The boundary slides just far enough to put your target on the class-0 side, and the rest of the space barely changes, which is why the other 449 test points are almost all still correct.

What to check on a test

Clean label poisoning hides from the two checks people trust, the label audit and the accuracy score, so you test for it differently.

CLEAN LABEL SURFACE
Who can edit training features, not just labels?
Are feature values versioned and hashed?
Any samples far from their own class centre?
Points oddly close to another class boundary?
Does accuracy hide a single flipped instance?
Are sensitive classes spot-tested on known inputs?
Feature-space outliers and per-instance tests, because the labels and the accuracy both look fine.

The finding to write up is a training pipeline where feature values can be edited without the same controls as labels, and where nothing tests the deployed model on known-good instances of the classes that matter. If an attacker can move feature rows and you only guard labels and watch overall accuracy, a targeted flip walks straight through. It's the same integrity gap the pipeline attacks exploit, one stage deeper. The AI pentest notes carry the full checklist.

Fix it, and detect it

Clean label attacks are harder to spot than label flips, so detection leans on the features themselves.

Guard the features like the labels. Whatever controls the training data has for labels, write access, versioning, hashing, apply them to feature values too. Most defences focus on labels and leave the features editable, which is the gap this attack lives in.

Hunt feature-space outliers. Look for training points that sit far from their own class centre, or unusually close to another class's boundary. The 5 perturbed neighbours in the lab are class-0 points sitting inside class 1, which stands out if you measure distance to class centroids rather than eyeballing labels.

Spot-test the classes that matter. Keep a small set of known-good, trusted instances of your sensitive classes and predict them after every retrain. A clean label attack aims at specific inputs, so testing specific inputs is how you catch it, where an aggregate accuracy score never will.

Watch the boundary between versions. Compare the decision boundary or model weights across retrains. A small, localised warp near one region, with overall accuracy unchanged, is the signature of a targeted clean-label push rather than normal drift.

Common mistakes

Three things trip people up building this.

Epsilon too small to cross. If epsilon_cross is too low, the perturbed neighbours don't actually reach the class-1 side and the boundary won't move enough. Check that each neighbour's f01 flips from positive to negative after the push; if it doesn't, raise epsilon (the lab used 0.25, the challenge needed 0.4).

Perturbing the wrong class. To flip a class-1 target towards class 0, you move class-0 neighbours into the class-1 region, and you don't touch class-1 points. Get the source and destination classes backwards and the boundary moves the wrong way, or not at all.

Editing the target itself. The target must stay untouched, that's the point: its features never change, only the model's opinion of them does. If your neighbour search accidentally includes the target, exclude it, or you're no longer running a clean-label attack on that instance.

Quick recap

  • A clean label attack changes the features of a few samples and keeps their real labels, so a label audit sees nothing wrong.
  • It's targeted: in the lab it flipped one chosen instance (index 373) while overall accuracy fell just 0.0022, from 0.9600 to 0.9578.
  • You pick a target already near the boundary (f01 = -0.0493), then push a few opposite-class neighbours across it, keeping their labels.
  • The retrained model warps its boundary to fit the odd points, and that small shift is enough to swallow the borderline target.
  • Detect it with feature-space outlier checks, per-instance spot-tests of sensitive classes, and boundary comparison between retrains, because accuracy won't show it.
  • Framework-wise it's OWASP LLM04:2025 Data and Model Poisoning (the old LLM03 name on the 2023 list).

Go deeper

The idea that a targeted poison can keep clean labels and preserve overall accuracy comes from the 2018 paper Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks by Shafahi and colleagues. They show it on real image classifiers, where a single poisoned image with a correct label can flip one target, and a "watermarking" trick with around 50 poisoned samples makes it reliable for end-to-end training. Reading it after this lab is the fastest way to see how the toy boundary shift scales to a real network.

If you're testing a training pipeline for this, or trying it in your own lab, I'm always happy to compare notes on LinkedIn.

References

FAQ

What is a clean label attack?

It's a targeted data poisoning attack that changes the features of a few training samples while keeping their original, correct-looking labels. The goal is to make one specific input get misclassified at inference time, without dropping overall accuracy, so the poisoned data looks normal on inspection.

How is a clean label attack different from label flipping?

Label flipping changes the labels and leaves the features alone; clean label does the opposite, changing features and keeping labels. Label flipping usually degrades a whole class, while a clean label attack aims at one chosen instance and keeps the model's overall accuracy almost unchanged.

Why are clean label attacks hard to detect?

Every label is still correct and overall accuracy barely moves, so both a label audit and an accuracy check pass. In a lab the accuracy dropped only 0.0022, from 0.9600 to 0.9578, while one specific point flipped. There's no obvious signal unless you look for feature outliers near the boundary.

How do you defend against clean label poisoning?

Control write access to training features, version and hash the dataset, and look for samples that sit far from their own class in feature space or unusually close to another class's boundary. Test the deployed model on known-good instances of sensitive classes to catch a targeted flip.