White Paper

Person Face Classification Using AuraFace Model and SVM

An AI-powered face recognition proof of concept comparing fine-tuned deep learning against an embedding-based pipeline for recognising individuals from a small, real-world dataset.

Technology AI
Technology Category
Jul 2026 Published
Abdul Basith P M Lead author

Executive Summary

The customer requires a solution capable of identifying and recognising individuals from images using a custom dataset maintained within the organization. The solution is intended to support scenarios where an individual’s identity needs to be determined from captured images without relying on manually labelled or pre-defined image metadata.

To address this requirement, a Proof of Concept (POC) was developed to evaluate the feasibility of building a person face recognition solution using deep learning-based face embeddings and machine learning classification techniques. The POC uses AuraFace to generate facial embeddings from images and an SVM classifier to identify individuals based on those embeddings.

The primary objective of the POC is to validate whether the proposed approach can accurately recognise individuals from the organization's custom dataset and provide a foundation for further development of a production-ready solution.

For production deployment, the solution would require appropriate privacy and compliance controls, including consent where required, data retention and deletion policies, secure storage and access controls, and assessment of applicable regulatory and organizational requirements. These considerations are outside the scope of the current POC but must be addressed for production readiness.

Problem in Context

Developing a face classification model with the person dataset has some technical challenges. Unlike general image classification, facial recognition requires the model to learn subtle differences between facial structures.

The person dataset available here contains a limited number of images per person, especially when compared to datasets on Kaggle, making direct fine-tuning prone to overfitting. Adding to the challenge, the image quality is often poor, and many photos feature groups with multiple people rather than isolated, clear portraits.

Technical Deep Dive

The face recognition system was evaluated using two fundamentally different approaches: transfer learning with ResNet34 and an embedding-based recognition pipeline using AuraFace and SVM. Although both methods achieved high training accuracy, their behavior during real-world testing differed significantly.

Preprocessing

Standard preprocessing techniques, including resizing, normalization, rotation, and data augmentation, were applied before training.
Reference image from the dataset: a group photograph of four colleagues in matching NKORR polo shirts standing outdoors in front of palm trees.
Reference Image
The images shown below have been preprocessed using standard transformations, including resizing, random rotation, random cropping, normalization, and other data augmentation techniques.
Four preprocessed face crops extracted from the group photograph, each labelled with the name of the person: Nithin, Ananthu, Hari Krishnan B S, and Ajay.
After Preprocessing

Fine-Tuned ResNet34

A pretrained ResNet34 model was fine-tuned by replacing the final classification layer to match the number of persons in the dataset. The model achieved over 90% training and validation accuracy. However, testing showed that predictions became unreliable when images contained excessive background information or inconsistent framing. This indicated that the network had learned background context alongside facial features, a common issue when training with relatively small datasets. Cropping images to include only the face reduced this problem and improved prediction consistency, but it also made the model dependent on consistent preprocessing during inference.

AuraFace + SVM Pipeline

The second approach separates feature extraction from classification. AuraFace generates high-dimensional facial embeddings that capture distinctive facial characteristics while remaining less sensitive to background variations. These embeddings are then used to train an SVM classifier.
Architecture diagram with two panels. The training phase on the backend runs from a labelled dataset through the AuraFace embedding model to SVM training, saving a trained SVM model. The inference phase on the device runs from camera capture through the AuraFace embedding model to SVM prediction and an identity-verified result, with the trained SVM model deployed across to it.
Architecture
This architecture offers several advantages:
  • Robust feature extraction using a pretrained model
  • Lightweight and fast classifier training
  • Simple retraining by updating embeddings and the SVM only
  • Better generalization on limited datasets

Tradeoffs

Aspect Fine-Tuned ResNet34 AuraFace + SVM
TrainingEnd-to-end fine-tuningFixed embeddings + SVM
Background SensitivityHigherLower
RetrainingNeural network retrainingSVM retraining only
Dataset RequirementLarger dataset preferredGood for smaller datasets
Real-world RobustnessModerateHigh

Model Results

Image 1: a close-cropped outdoor portrait with a palm tree background.
Image 1
Image 2: a full-length outdoor photograph in which the face occupies only a small part of the frame.
Image 2
Image 3: an indoor portrait against a plain background.
Image 3
Image Model Confidence Correctness
Image 1Fine-Tuned ResNet34Hari Krishnan B S (99.96 %)✅
Image 1AuraFace + SVMHari Krishnan B S (63 %)✅
Image 2Fine-Tuned ResNet34NAJIYA P P K (25.01 %)❌
Image 2AuraFace + SVMHari Krishnan B S (68 %)✅
Image 3Fine-Tuned ResNet34Hari Krishnan B S (99.93 %)✅
Image 3AuraFace + SVMHari Krishnan B S (76 %)✅
Note: The confidence scores represent the model's estimated probability for the predicted class and should not be interpreted as a direct measure of model accuracy. Although the fine-tuned ResNet34 produces near-100% confidence scores, it also generates incorrect predictions with high confidence, indicating overconfident behaviour. AuraFace + SVM produces more conservative confidence scores while achieving more reliable recognition on the evaluated samples.

NKORR’s Approach

At NKORR, we approach face recognition as a feature representation problem rather than solely a classification task. Instead of retraining deep neural networks for every deployment, we leverage pretrained embedding models to generate robust facial representations and train lightweight classifiers for identity recognition. Our approach is guided by four principles:
  • Leverage pretrained knowledge to reduce training time and improve generalization.
  • Separate feature extraction from classification to simplify updates and maintenance.
  • Design for scalability, allowing new persons to be added with minimal retraining.
  • Validate under real-world conditions, focusing on reliability across varying backgrounds, lighting conditions, and image quality rather than benchmark accuracy alone.
This modular design results in a solution that is easier to maintain, scalable as datasets grow, and better suited for production environments.

Recommendations

  • Evaluate models using real-world images rather than relying solely on validation accuracy.
  • Use embedding-based approaches when working with limited datasets or when identity updates occur frequently.
  • Apply consistent face detection, alignment, and cropping during both training and inference.
  • Separate feature extraction from classification to reduce retraining effort and improve maintainability.
  • Use confidence thresholds and periodic model evaluation to maintain reliable recognition performance in production.
Read the original white paper This page carries the full content of the paper. Download the PDF for the original typeset version.
Download PDF

Authors & Contributors

Abdul Basith P M Abdul Basith P M Software Engineer
Vishak Kurup Vishak Kurup Technical Architect

References

#SourceNotes
AuraFace Model - https://huggingface.co/fal/AuraFace-v1HuggingFaceModels in HuggingFace
Practical Deep Learning for Coders - https://course.fast.ai/Jeremy HowardFastai
Introduction to Pytorch - https://docs.pytorch.org/docs/2.13/index.htmlPytorchLearn Pytorch

Work with us

Need to recognise people
from a small, messy photo dataset?

Tell us about your project. We'll tell you honestly whether and how we can help.