AIPractix Benchmarks

AI Models,
Measured in the Real World.

We measure the models behind our products across accuracy, operating thresholds, latency and real deployment conditions — so we know not only what works, but where it works.

Our principleMeasure honestly.
Compare globally.
Keep improving.

Public results are paired with their evaluation source and protocol wherever possible.

Independent evaluation · Face Recognition
NIST FRTE 1:1

AttFace, evaluated beyond our own lab.

AIPractix submitted aipractix-000 to the U.S. National Institute of Standards and Technology Face Recognition Technology Evaluation (FRTE) 1:1 track. NIST evaluates submitted algorithms under common protocols across multiple face verification datasets.

For us, independent evaluation is not a badge. It is part of how we build: measure against the same conditions as developers around the world, learn from the result, and improve.

View AIPractix NIST FRTE 1:1 submission results
Reported verification result*99.35%
Evaluation
NIST FRTE 1:1
Algorithm
aipractix-000
Submitted
May 12, 2026
Developer
AIPractix

* Face verification performance depends on the dataset and operating threshold. NIST FRTE reports false non-match rate (FNMR) at specified false match rates (FMR) across several datasets. This figure should not be interpreted as a universal accuracy for every deployment. See the official NIST results for evaluation context.

Model evaluation

What we measure.

A benchmark is useful only when its task, data, metric and environment are clear. We publish model-level results progressively as evaluation protocols and versions are locked.

AttFaceThird-party evaluated

Face Recognition

Verification and identification across identity workflows, with accuracy considered together with threshold selection and runtime performance.

TAR / FNMRFMRLatencyTemplate size
AttFaceInternal evaluation

Anti-Spoofing & Liveness

Presentation-attack detection evaluated separately from recognition across print, replay and other spoof scenarios.

ACERAPCERBPCERLatency
AttFaceInternal evaluation

Detection, Quality & Pose

Models that decide whether a face is present, usable and suitable for downstream recognition or monitoring.

RecallQualityPose errorRuntime
FeatSearchEvaluation in progress

Multimodal Retrieval

Image, semantic and similarity retrieval measured around relevance as well as the latency of searching practical index sizes.

Recall@KPrecision@KmAPSearch latency
Evaluation discipline

A number needs context.

Every result we publish should answer the same five questions, whether it came from a public benchmark, our own lab or a production-oriented test.

01What was tested?Exact task, model and version.
02On what data?Public, internal or deployment-oriented dataset.
03Which metric?Task-appropriate accuracy or error measure.
04On what hardware?CPU, GPU, mobile or edge environment.
05Under what conditions?Threshold, resolution, scale and relevant constraints.
Production AI

Accuracy is only part of the system.

A model with the highest score is not automatically the best production system. Real applications also depend on latency, robustness, scale, privacy and where the technology needs to run.

AccuracyLatencyScaleRobustnessPrivacyDeployment

From benchmark to real-world AI.

Explore AttFace and FeatSearch, or tell us about the workflow you need to build. We can help choose the right model, operating point and deployment architecture.

Discuss your project