How computers learned to tell nearly identical coffee beans apart
Researchers trained artificial intelligence models on roughly 190,000 images, achieving over 99% accuracy when identifying roasted coffee varieties.
Reading level
The full story with proper science words explained.

Heads up: this study is a preprint, which means other scientists haven’t finished checking it yet.
Spotting the tiny differences
Could you spot the difference between two roasted coffee beans? To human eyes, roasted beans look like identical little brown pebbles. Yet getting automated coffee bean classification right matters when verifying what buyers are paying for.
Researchers set up a custom rig in a controlled space to photograph beans from 15 suppliers. They captured 8,417 original high-resolution pictures. After processing, this grew into about 190,000 images resized to standard dimensions, such as 299 by 299 pixels.
Testing the digital detectives
The team trained several artificial intelligence systems to spot subtle varieties. They tested deep learning models known as a CNN (convolutional neural network), including designs called ResNet and EfficientNet-B0. They also tried traditional machine learning methods like Random Forest and XGBoost, alongside a hybrid mix.
Think of the software like facial recognition for twins. Even when two faces look the same at first glance, the computer maps tiny contours to tell them apart. However, while human faces move, beans stay frozen in place under fixed light.
To measure success fairly, researchers used stratified group cross-validation, which splits data into balanced test groups. The standout winner was EfficientNet-B0. It achieved an average accuracy of 99.86% when sorting four main bean categories. Even when separating eight closely related Arabica subtypes, the model scored 99.73% accuracy. Among non-deep algorithms, XGBoost reached 97.98%.
What we still do not know
These near-perfect numbers come with limits. The pictures were captured under strictly controlled lighting with beans from only 15 suppliers. Roasters around the globe might produce colours and surfaces that differ from this batch.
We also do not know how these models would cope with everyday smartphone snapshots or high-speed conveyor belts. Testing unroasted green beans or mixed blends remains unexplored for now.
Science words
- CNN
- A type of artificial intelligence designed specifically to analyse visual imagery.
- EfficientNet-B0
- A lightweight type of neural network designed to process images quickly and accurately.
- XGBoost
- A machine learning method that builds decision trees step by step to correct earlier mistakes.
- Stratified group cross-validation
- A testing method that splits data into balanced subsets to make sure performance results are reliable.
Check it yourself
This story is based on a real research paper in Scientific Publication by Hastürk, Sümer. We write with AI help and check it against the paper, but the original is the final word.