I am the founder of Metric AI Lab, a research lab
in Yerevan, Armenia. We work on Physical AI, and most of my time goes to failure detection
for robot policies: getting a policy to signal that it has gone wrong, early enough that
something can be done about it. The lab also builds open retrieval and embedding models for
low-resource languages, several of which are still in wide use. I hold a PhD in economics and
taught data science at the American University of Armenia.
Sep 2026We released
ArmBench-ASR, the first
benchmark for Armenian speech recognition. Version 1.0 scores 34 open and closed systems
across read speech, poetry, film dialogue and narrated news. Closed systems take the top
ten places; the strongest open model lands eleventh.
Sep 2026FailBench is on arXiv, a benchmark for
robot failure detection built from 2,197 manipulation attempts across 14 public sources.
The best of thirteen VLM-based detectors reaches only 0.77 mean balanced accuracy, and
models fine-tuned specifically for failure detection consistently underperform the
general-purpose models they were fine-tuned from.
Sep 2026Our work on Uzbek legal RAG was accepted to the
EMNLP 2026 Industry Track.
We release UTE-1, a state-of-the-art open Uzbek text embedder, along with a retrieval
benchmark and an end-to-end legal QA benchmark.
Aug 2026Our audit of Physical AI benchmark redundancy is on
arXiv. Four of twelve benchmarks carry 78.5%
of the signal, and collapsing the two substitute pairs moves 22 of 51 models by three or
more ranking places.
Apr 2026We released ArmBench-LLM 1.0,
a benchmark comparing frontier and open models on Armenian knowledge and generation tasks.
Mar 2026Less is More was presented at
LoResLM at EACL 2026 in Rabat, alongside the release of
ATE-2, our Armenian text embedding
models and benchmark.
OngoingThe lab is hiring research engineers for pre-training and post-training of multimodal
foundation models. Write to hr@metric.am.
Research
My work in Physical AI runs along three connected threads. Nearly all of my time goes to the
first, and the third is where my own training is most useful.
Failure detection
Robot policies act confidently and fail silently. A policy that cannot signal its own failure
cannot be trusted to run unsupervised, which is the largest obstacle between a working demo and
a deployed system. The tools currently used to make that judgment are not good enough: on
FailBench, the best of thirteen vision-language detectors gets roughly one judgment in four
wrong, and falls to chance on contact-rich assembly. We work on detecting failure at runtime,
from the policy's own internal state rather than from an external judge, and early enough that
intervention is still possible.
World models
Action-conditioned latent prediction is the substrate that runtime detection runs on: a model
that can anticipate the next state can also notice when the world stops matching what it
expected. We collect and release real and simulated manipulation data to train these models.
Evaluation and benchmark science
My training is in econometrics, so I care about whether the numbers the field reports mean what
they are taken to mean. Physical AI benchmarks have multiplied faster than anyone has checked
whether they measure different things. Our audit builds one comparable matrix across 51 models
and 12 benchmarks, shows quantitative evidence of redundancy, and selects a four-benchmark
subset that retains most of the signal.
Retrieval and low-resource languages
The lab's earlier line of work, still active. Open visual document retrievers, text embedding
models, and production retrieval systems for languages where the binding constraint is that the
data does not exist at the scale the standard recipes assume: Armenian, Uzbek and others.
Selected projects and publications
Physical AI
FailBench: How Reliable are VLMs at Judging Robot Task Success? Preprint, 2026 Z. Navasardyan, T. Danielyan, H. Davtyan arXiv ·
pdf ·
project
A Statistical Audit of Physical AI Benchmark Redundancy Preprint, 2026 Z. Navasardyan, H. Davtyan arXiv ·
pdf ·
project ·
code
Robot manipulation datasets for failure detection Open datasets, 2026 Metric AI Lab
Paired simulated and real trajectories on the SO-101 arm.
datasets
Retrieval and low-resource languages
Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever
Fine-Tuning EMNLP 2026, Industry Track T. Danielyan, M. Avetisyan, H. Davtyan project ·
UTE-1 model
Less is More: Adapting Text Embeddings for Low-Resource Languages with
Small Scale Noisy Synthetic Data LoResLM at EACL 2026 Z. Navasardyan, S. Bughdaryan, B. Minasyan, H. Davtyan ACL Anthology ·
arXiv ·
project
ArmBench-ASR: Benchmarking Speech-to-Text Models on Armenian Benchmark and leaderboard, 2026 Metric AI Lab
The first benchmark for Armenian speech recognition: 34 open and closed systems across five
speech domains, from crowdsourced read speech to film dialogue.
writeup ·
leaderboard
ArmBench-LLM Benchmark and leaderboard, 2026 Metric AI Lab
Evaluation of frontier and open models on Armenian knowledge, generation and reasoning,
with a cost-versus-accuracy report.
writeup ·
leaderboard
ATE-2: Armenian text embeddings and the ArmBench-TextEmbed benchmark Open models, data and benchmark, 2026 Metric AI Lab writeup ·
models ·
leaderboard
ColQwen2.5-multilingual: visual document retrievers at 3B and 7B Open models, 2025 Metric AI Lab
Multilingual late-interaction retrievers for visually rich documents.
models
ColQwenStella-2b-multilingual Open models, 2025 Metric AI Lab model
Teaching and background
I hold a PhD in economics, and studied at University College London and the International School
of Economics at Tbilisi State University. I taught data science at the American University of
Armenia and, through open courses and course materials, to several thousand students beyond it.
I have also worked as a contracted expert for UN agencies on statistics and applied econometrics.
Metric AI Lab grew from one person into a team that has delivered more than a hundred enterprise
AI projects. That work funds the research.