Fintech AI · Explainer
Fintech AI

What AI Can and Cannot Read From a Face

Identity verification from a face is standardised, independently tested and required by regulators. Inferring emotion, stress or deception from the same video is contested science with no independent benchmark this site could find and, in some settings, is now prohibited. The evidence for both, with sources.

Verified September 2026Free · No signupOfficial sources only
BeginnerStart here. No prior knowledge assumed.

The short version

Point a camera at someone’s face and software can do two very different kinds of thing.

It can check who they are — is this the same face as the passport, is there a live human here rather than a photograph. That works. It is measured by national standards bodies, it has published error rates, and regulators require it.

It can also claim to tell you what they are feeling — are they stressed, confident, nervous, lying. That is a different proposition with a much weaker evidence base, no independent benchmark this site could find against which to check a vendor’s accuracy claim, and in some settings it is now illegal.

These two get sold together, in the same demo, by the same vendor, on the same screen. This page is about telling them apart.

In one paragraph

The short version. Identity verification from a face is a solved, standardised, independently-measured technology with known and fixable weaknesses. Emotion, stress and deception inference from a face is not — and the EU has written the scientific critique of it into binding law. If a vendor bundles the second into a quote for the first, that is the moment to slow down.

The two columns

Here is the whole page in one table.

Column A — worksColumn B — contested
What it doesIs this the right person? Is this a live human?What is this person feeling or intending?
ExamplesFace matching against an ID, liveness and spoof detection, document authenticityEmotion recognition, stress detection, confidence scoring, deception detection
Independent standardISO/IEC 30107-3 for spoof testing; NIST evaluations for matching accuracyNone found
Published error ratesYes — by algorithm, by demographic groupNo agreed metric to publish
Indian regulatorRBI requires it for video KYCNot required, not mentioned
EU AI ActOne-to-one verification carved out of the biometric high-risk listProhibited at work and in education; high-risk everywhere else
Where the line is drawn

That split is not a framing this page invented. It is the structure of the statute. Annex III of the EU AI Act lists high-risk biometric systems in a single numbered point: sub-paragraph (a) carves out one-to-one verification “the sole purpose of which is to confirm that a specific natural person is the person he or she claims to be”, and sub-paragraph (c) pulls in “AI systems intended to be used for emotion recognition” [21]. Two adjacent sub-paragraphs, opposite treatment.

Why this page exists

Because the second column is being sold into Indian fintech right now, on top of products that do the first column perfectly well.

The pitch is reasonable-sounding: you are already recording a video call for KYC, the customer is already on camera, so why not also score whether they seem nervous and flag the risky ones? No extra hardware, no extra friction, one more line on the invoice.

The problem is that every part of that sentence is true except the part where it works.

IntermediateBuild it. Pipelines, tools and working code.

Column A — what actually works

Start with the half that does work, because taking it seriously is what earns the right to be sceptical about the other half.

Liveness and spoof detection

ISO/IEC 30107-3:2023 is the international standard for testing whether a system can tell a real face from a presented fake. Its published scope covers the performance assessment and reporting of presentation attack detection [16]. The standard itself — which defines the two error rates the industry quotes, APCER for attacks wrongly accepted and BPCER for genuine users wrongly rejected — sits behind ISO’s paywall, so this page works from the catalogue scope and from iBeta’s published methodology below rather than from the text.

iBeta is the best-known independent lab testing against it. Its published levels are worth understanding precisely, because vendors quote them as though they meant “spoof-proof” and they do not:

Level 1Level 2
Time per attack species8 hours2–4 days
Attacker expertise assumedNoneHas done at least one prior PAD test
Materials budget cap$30$300
Attack success allowed0%1%
Genuine users it may rejectUp to 15%

Source: iBeta’s own published methodology [17].

So a Level 1 certificate means: resistant to a thirty-dollar attack, by someone with no relevant expertise, given one working day. That is a real, testable, meaningful claim. It is also much narrower than the badge implies — and note the last row, which allows a compliant system to reject up to fifteen percent of your genuine customers.

Watch out

The gap that matters most in 2026: injection attacks are outside the standard. ISO/IEC 30107-3 explicitly scopes out attacks occurring anywhere other than “at the biometric capture device during presentation” [16]. An injection attack does not hold a screen up to the camera — it replaces the camera feed entirely with a virtual driver and pipes synthetic video straight into your application. Every pixel-level check passes, because the pixels are perfect. An “ISO 30107-3 certified” badge says nothing whatsoever about deepfake injection resistance. NIST's own PAD evaluation excludes digital injection attacks too [20]. Ask vendors about injection separately, in writing.

Face matching

NIST runs the long-standing independent evaluation, now split into FRTE for recognition and FATE for face analysis [18]. Its demographic study remains the reference point: 189 algorithms from 99 developers, tested against 18.27 million images of 8.49 million people [19].

The findings, in NIST’s words:

  • “False positive rates often vary by factors of 10 to beyond 100 times” across demographic groups.
  • “False positive rates are highest in West and East African and East Asian people, and lowest in Eastern European individuals.”
  • “…false positives to be between 2 and 5 times higher in women than men” — with the multiple varying by algorithm, country of origin and age.
  • “…elevated false positives in the elderly and in children.”
  • But false negatives “tend to be more algorithm-specific, and vary often by factors below 3”.

That is a serious problem and this page is not minimising it. But read the next finding carefully, because it is the one most often left out of the summary:

“A number of algorithms developed in China give low false positive rates on East Asian faces, and sometimes these are lower than those with Caucasian faces.” [19]

The differential is a property of what the algorithm was trained on, not of faces. It is an engineering failure, which means it is fixable, and NIST’s finding of a “wide range in accuracy across developers, with the most accurate algorithms producing many fewer errors” means choosing your vendor is a real lever.

Hold on to that, because it is precisely the lever that does not exist in Column B.

What India actually requires

The RBI’s KYC Master Direction requires that a video KYC application have “components with face liveness / spoof detection as well as face matching technology with high degree of accuracy”, and adds that “appropriate artificial intelligence (AI) technology can be used to ensure that the V-CIP is robust” [24].

Both halves of Column A, named explicitly. Note what that AI sentence is actually about: the robustness of the identification process. It is the sentence most likely to be quoted at you out of context.

Column B — reading emotion from a face

Now the other column.

The review that defines the field

In 2019 five researchers — Lisa Feldman Barrett, Ralph Adolphs, Stacy Marsella, Aleix Martinez and Seth Pollak — published a 68-page systematic review in Psychological Science in the Public Interest examining whether you can infer emotion from facial movement [1].

Be fair to what it found, because the nuanced version is stronger than the caricature. It did not conclude that faces tell you nothing. From the abstract:

“The available scientific evidence suggests that people do sometimes smile when happy, frown when sad, scowl when angry, and so on, as proposed by the common view, more than what would be expected by chance. Yet how people communicate anger, disgust, fear, happiness, sadness, and surprise varies substantially across cultures, situations, and even across people within a single situation. Furthermore, similar configurations of facial movements variably express instances of more than one emotion category.” [1]

Better than chance. Not reliable, not specific, not generalisable. Their verdict:

“…hypothesized facial configurations are not observed reliably or specifically enough to justify using them to infer a person’s emotional state, whether in the lab or in everyday life.” [1]

The numbers behind it: an average correlation of r = .31 between the intensity of a facial configuration and the corresponding emotion, which the authors characterise as weak evidence of reliability; and a hypothesised expression actually appearing during the matching emotional event about 22% of the time [1].

The missing number

And one absence that should shape how you read every vendor datasheet. On specificity, the review reports that no overall assessment could be made — “because most published studies do not report the false-positive rate” [1]. That is the same gap you will find in commercial accuracy claims. A system that flags stress correctly 90% of the time is worthless if it also flags calm people 40% of the time, and the false-positive rate is the number nobody volunteers.

The other side of the argument

There is a real scientific dispute here and this page is more useful for naming it. The same journal issue carried a companion article by Alan Cowen, Disa Sauter, Jessica Tracy and Dacher Keltner arguing that the six-basic-emotion framework is the wrong target, and that a higher-dimensional taxonomy of emotional expression does show systematic structure [2].

That is a genuine challenge to Barrett’s framing. Note, though, what it is not: it is a disagreement about how emotion should be modelled, not a demonstration that a camera can tell you whether a loan applicant is being honest.

A 2021 meta-analysis in Emotion came down on Barrett’s side, concluding that “what are commonly known as the six classic basic emotions do not reliably co-occur with their predicted facial signal”, with co-occurrence effect sizes of 0.13 for a whole predicted expression and 0.23 for whole-or-partial [3]. A published comment in the same journal disputes that analysis, under the title “Emotions do reliably co-occur with predicted facial signals” [4] — which states its position plainly enough. This page goes no further than the title, because this site could not retrieve the full text to check the argument behind it.

What happens when you test a commercial system

A 2026 cross-corpus evaluation ran a commercial automatic facial coding system against five spontaneous-emotion databases. Overall accuracy was 34.4% against a 20% chance baseline [5].

The breakdown is more interesting than the headline. Happiness: 94%. Surprise and disgust: around 24%. Fear and sadness: 14–16%. And the failure mode is systematic — the system misclassified 44.9% of surprise, 61.7% of fear and 58.7% of disgust as “happiness” [5].

A system that reads most fear as happiness is not a weak emotion detector. It is substantially a smile detector with a confident user interface.

Detecting deception

Deception detection deserves its own section because it is the claim most likely to be made to a fraud team, and it has the weakest evidence of anything on this page.

Humans are barely above chance

The definitive meta-analysis synthesised 206 documents covering 24,483 judges. People correctly classify lies and truths 54% of the time, against a 50% baseline — and that breaks down as 47% of lies correctly identified against 61% of truths [6].

Read that asymmetry again. The above-chance performance is largely a bias toward believing people, which pays off because most statements in these studies are true. And crucially for anything video-based, judges were more accurate judging audible than visible lies [6]. The face is the weaker channel, not the stronger one.

The cues themselves are weak, and may be weaker than published

The largest review of behavioural cues to deception examined 158 cues across 120 independent samples and found that “many behaviors showed no discernible links, or only weak links, to deceit” [7].

A 2019 analysis went further, using simulations to show that reported effect sizes in this literature are likely inflated by publication bias and low power. Its conclusion is the strongest honest statement available on the subject:

“The informational value of the present deception literature is quite low, such that it is not possible to determine whether any given effect is real or a false positive.” [8]

Do machines do better? The honest answer is: nobody has shown it

A 2023 systematic review in PLOS ONE covered 81 peer-reviewed machine-learning deception studies. Reported accuracies span 51% to 100%, with 19 studies above 90% — a spread the authors attribute to incomparable experimental conditions rather than genuine progress [9].

The findings that matter:

  • 70% of studies used mock deception — people instructed to lie in a lab — rather than real-world data.
  • The most-reused real-life dataset contains 121 video samples. Some real-life datasets had as few as six.
  • The authors’ own verdict: “Industry application looks premature at present. Most experiments are based on mock data.” [9]

So when a vendor shows you 95% deception-detection accuracy, the overwhelmingly likely explanation is an in-domain result on acted lies, from a corpus of perhaps a hundred videos, tested on data resembling its training set. It is not evidence about your customers.

Pupil dilation and stress

This one is worth care, because unlike deception detection there is real science underneath — which is exactly what makes the overreach persuasive.

What is genuinely established

Pupils do dilate with cognitive load and arousal. Kahneman and Beatty demonstrated it in Science in 1966 [10], and it has held up for sixty years. The modern authoritative review confirms the effect is real and not trivial in size [11].

Why that does not let you read a mind

From that same review:

“Anything that somehow activates the mind, or anything that increases the mind’s ‘processing load’… also causes the pupil to dilate.” [11]

And more bluntly:

“Attempts to link specific cognitive processes to the [pupillary response] (not mediated by processing load) are therefore, in my view, doomed to fail.” [11]

Pupil dilation is a many-to-one signal. Effort, interest, surprise, arousal, anxiety, a difficult question, an unfamiliar interface, caffeine — all produce it. Nothing in sixty years of research makes dilation diagnostic of deception rather than of thinking.

The lighting problem is not a caveat, it is the whole thing

The pupil changes the light entering the eye by a factor of roughly 16. Cognitive effects on pupil size are, in the review’s words, “generally modest” — from under 1% to about 5% [11].

The confound is an order of magnitude larger than the signal. In a lab with fixed luminance you can subtract it. On a phone, with ambient light changing, the user moving relative to a window, and the screen itself changing brightness as the interface changes, you cannot — and worst of all, the screen brightness changes at the same moments the questions do. The confound is correlated with your independent variable.

Even in ideal lab conditions it does not work per-person

A 2022 study used research-grade eye tracking in a controlled paradigm to detect concealed information. Result: “while most participants showed this effect qualitatively, it was not statistically significant for most participants individually”, and the authors concluded that further development is needed for “a valid and reliable concealed identity information detector at the individual level” [12].

A group-level effect that does not reach significance per person is useless for a product that must make a decision about one applicant.

Watch out

And on an Indian user base, the measurement often fails before the inference does. Published research on smartphone pupillometry notes that “the melanin in the iris absorbs light such that it is difficult to distinguish the pupil aperture that appears black”, and that prior work “either exclude[s] individuals with dark iris colors from the analysis or report[s] poor performance” [13]. A far-red workaround produced a 624% contrast improvement for dark-brown-eyed participants — which quantifies how badly the ordinary visible-spectrum approach fails. On standard phone cameras, across most Indian users, a pupillometry feature is largely measuring segmentation noise.

Heart rate from video

Reading heart rate from ordinary video — remote photoplethysmography — increasingly appears in video KYC products as a stress signal. Here the honest answer splits in two.

Heart rate: broadly real. Heart rate variability: not under these conditions

A 2026 scoping review asked precisely the question a fintech needs answered, and its finding is quotable:

“rPPG supports robust non-contact heart-rate estimation, while HRV/PRV reliability degrades under head motion, illumination variability, and short analysis windows.” [14]

That distinction is the entire story, because every stress claim rests on variability, not on rate. The review’s stated minimum conditions for a defensible inference: analysis windows of at least 120 seconds for time-domain HRV, with windows under 60 seconds restricted to heart rate and coarse trends only; signal-quality monitoring with explicit abstention on poor segments; and validation against ECG or high-quality contact sensors [14].

Now describe a video KYC call: around a minute, the subject talking, head moving, lighting whatever the room offers, video compressed for transmission. That sits squarely inside the review’s own explicit do-not-infer conditions.

The skin-tone differential is published and large

An analysis of 100 rPPG studies found datasets skewed heavily toward participants of European or East Asian descent, and collated the performance gap [15]:

StudyLighter skinDarker skin
Dasari et al., traditional methods5.2 bpm error14.1 bpm
Dasari et al., deep learning6.0 bpm error9.5 bpm
Nowara et al.4.23 bpm error13.58 bpm

Read those rows separately rather than averaging them. The gap is about 2.7× and 3.2× for the traditional methods, but 1.6× for the deep-learning approach — which is what most vendors now ship. The penalty is real and large, and it is also narrowing with better models.

Even at 1.6×, though, it lands on the measurement everything downstream depends on. For an Indian customer base that is not a fairness footnote; it is a core accuracy problem that pushes the derived stress score toward noise for a large share of your users.

AdvancedShip it. Failure modes, thresholds and evidence.

The law: the EU AI Act

This is where the science and the law stop being two separate arguments.

What the Act covers

The EU AI Act defines an emotion recognition system as one “for the purpose of identifying or inferring emotions or intentions of natural persons on the basis of their biometric data” [21].

Note “or intentions”. A product marketed as detecting intent to defraud is inside this definition, not outside it.

The prohibition

Article 5(1)(f) prohibits “the placing on the market, the putting into service for this specific purpose, or the use of AI systems to infer emotions of a natural person in the areas of workplace and education institutions, except where the use of the AI system is intended to be put in place or into the market for medical or safety reasons” [21].

Where it is not prohibited, it is high-risk. Annex III point 1(c) captures emotion recognition generally [21], and the Commission’s guidelines on prohibited practices address the same fallback [23] — a document this site could not open directly, so it is cited rather than quoted.

The recital that ties this page together

Recital 44

Recital 44 of the Act, verbatim: “There are serious concerns about the scientific basis of AI systems aiming to identify or infer emotions, particularly as expression of emotions vary considerably across cultures and situations, and even within a single individual. Among the key shortcomings of such systems are the limited reliability, the lack of specificity and the limited generalisability.” [21]

Compare that with the review in the section above. Barrett and colleagues set out four criteria a facial-inference claim must meet — reliability, specificity, generalizability and validity [1] — and found the evidence failed on the first three.

The EU legislature listed the same three shortcomings, in the same order, in the operative recital of a regulation carrying penalties of up to €35 million or 7% of worldwide annual turnover [21] — the highest tier in the entire Act, above GDPR’s 4%.

This was not a privacy decision that happens to align with the science. The regulator adopted the scientists’ reasoning explicitly.

What applies today, and what was delayed

The Act’s timeline was amended in July 2026 by the Digital Omnibus on AI, which deferred parts of the high-risk regime [22]. Reading that as “emotion recognition is fine until 2027” gets it backwards. As at this page’s check date:

ProvisionApplies fromStatus today
Article 5(1)(f) prohibition2 February 2025In force
Article 99 penalties2 August 2025In force
Article 50(3) transparency duty2 August 2026In force
Annex III high-risk obligations2 December 2027 (was 2 August 2026)Not yet
Annex I embedded high-risk2 August 2028 (was 2 August 2027)Not yet

The prohibition and the disclosure duty are live now. Only the compliance machinery moved.

Article 50(3) requires deployers of an emotion recognition system to inform the people exposed to it of its operation, and Article 50(5) requires that information at the latest at the time of first interaction [21].

Watch out

If you are building an AI interview tool, assume recruitment is inside the prohibition. Whether hiring counts as “workplace” for Article 5(1)(f) decides whether an emotion-inferring interview product is prohibited at 7% of global turnover or merely high-risk. The Commission's guidelines on prohibited practices address the scope of the term, and are reported to read it broadly — covering any physical or virtual space where work is performed, and extending to recruitment and hiring rather than only to existing employees. This site could not open that guidance document itself, so it is cited but not quoted here [23] — but the direction is clear enough that the safe assumption is prohibition rather than high-risk. Annex III point 4(a) does separately treat recruitment as high-risk [21]; that is the regime for recruitment tools which do not infer emotion, not a licence for ones that do. Take advice before building either.

The law: India

The RBI neither requires nor mentions it

The V-CIP provisions of the KYC Master Direction require liveness, spoof detection and face matching [24]. On emotion, affect, stress or deception, this site found no mention anywhere in the Direction.

That produces a cleaner argument than a prohibition would. In India, emotion inference in a KYC context is not required, not contemplated, and earns you no regulatory credit whatsoever. A vendor selling it as “RBI-compliant enhanced KYC” is selling an unrequested feature that adds legal exposure and returns nothing.

DPDP: the absence of a biometric category makes this harder, not easier

The Digital Personal Data Protection Act, 2023 has no “sensitive personal data” category and does not mention biometric data at all [25] — a deliberate break from both the GDPR and India’s own earlier SPDI Rules, whose statutory basis the Act repealed [25].

Founders sometimes read that as permission. It is not, and the reason is section 6(1): consent must be “free, specific, informed, unconditional and unambiguous” and “limited to such personal data as is necessary for such specified purpose” [25].

A customer consenting to identity verification has consented to a face match. Running affect or deception inference over the same video is a different purpose, and that data is not necessary for the verification purpose. On the plain words of the section, it needs its own specific, informed consent.

Write the consent notice first

A useful test. Write the consent line you would have to show: “We will analyse your facial expressions and voice to infer your emotional state and assess whether you may be being deceptive.” If that sentence is one you would not want on the screen, the feature is not one you should be shipping. The discomfort is the compliance analysis.

The DPDP Rules were notified in November 2025 with an 18-month phased compliance timeline, and require a separate, clear consent notice explaining the specific purpose. The headline penalty — up to ₹250 crore — attaches specifically to a failure to maintain reasonable security safeguards rather than to the consent duty. Significant Data Fiduciaries must carry out impact assessments and “follow stricter checks while using new or sensitive technologies” [26].

An emotion-inference layer is unambiguously a new or sensitive technology. Which means an impact assessment — into which, per everything above, there is no validity evidence to put.

The asymmetry, and what to ask

Here is the structural fact that should settle it.

For everything in Column A there is an independent body publishing error rates against an agreed metric. ISO/IEC 30107-3 defines APCER and BPCER. iBeta tests against it. NIST evaluates 189 face-matching algorithms against 18 million images and publishes the demographic breakdown.

For Column B, this site searched ISO and NIST and found no standard specifying test and reporting methodology for emotion recognition accuracy, no NIST evaluation programme, and no independent conformity-assessment body. What exists instead are academic benchmark datasets — research corpora with no agreed error metrics, no independent testing houses, no pass criteria, and no validated ground truth, since the labels are themselves human guesses about inner states from faces.

Stated carefully

This is a negative search result, so it is stated as one: no such standard or independent benchmark appears to exist. This site cannot prove none exists anywhere. But the asymmetry is not marginal — it is the difference between an industry with independent referees and one without any.

The questions to ask a vendor

Not “is your 94% accurate?” — you cannot check that. Ask instead:

  • What is the false-positive rate? The review in the section above found most published emotion studies do not report one [1]. Most datasheets do not either.
  • Who measured it, against what ground truth? If the answer is an internal test set with labels assigned by human annotators guessing from faces, the ceiling of the system is the accuracy of the guess.
  • Was it validated on real behaviour or acted behaviour? Seventy percent of the machine-learning deception literature used instructed lies in a lab [9].
  • What is the error rate broken down by skin tone? For anything camera-derived, on an Indian user base, this is the question that decides whether the feature works for your customers [15][13].
  • Show me the independent certificate. For liveness there is one to show. For emotion this site could find nowhere to get one — and a vendor who can prove otherwise has just answered the question.

What to do instead

Everything Column B claims to deliver — fraud signal, risk scoring — has better sources. Device and network signals, behavioural biometrics on the application itself, document forensics, cross-referencing against your own history with that customer, and injection-attack detection on the video path. All of them are measurable, none of them requires guessing at somebody’s inner state, and none of them carries a 7%-of-turnover prohibition.

Build Column A properly. It is harder than it looks, the injection-attack gap is real and current, and doing it well is worth more than any amount of inferred emotion.

Where to go next

Read next:

Sources

Every claim on this page traces to one of these: peer-reviewed research, an official regulatory text, or a standards body. Where this site could not retrieve a primary source, the page says so rather than paraphrasing from secondary coverage.

  1. researchBarrett, Adolphs, Marsella, Martinez & Pollak (2019), Psychological Science in the Public Interest 20(1):1–68 — “Emotional Expressions Reconsidered” — the four criteria, the r = .31 and 22% figures, the absence of reported false-positive rates, and the conclusion that facial configurations are not observed reliably or specifically enough to infer emotional state. https://journals.sagepub.com/doi/full/10.1177/1529100619832930
  2. researchCowen, Sauter, Tracy & Keltner (2019), Psychological Science in the Public Interest 20(1) — “Mapping the Passions” — the companion article arguing for a higher-dimensional taxonomy, published as a counterpoint in the same issue. https://journals.sagepub.com/doi/10.1177/1529100619850176
  3. researchDurán & Fernández-Dols (2021), Emotion — meta-analysis finding the six classic basic emotions do not reliably co-occur with their predicted facial signal; co-occurrence effect sizes of 0.13 and 0.23. https://pubmed.ncbi.nlm.nih.gov/34780241/
  4. research“Emotions do reliably co-occur with predicted facial signals: Comment on Durán and Fernández-Dols (2021)”, Emotion (2023) — a published dispute of that meta-analysis. Cited by title; this site could not retrieve the full text and does not characterise the argument behind it. https://pubmed.ncbi.nlm.nih.gov/37079837/
  5. researchCross-corpus evaluation of a commercial facial coding system, Electronics 15(4):849 (2026) — 34.4% overall accuracy against a 20% baseline, the per-category breakdown, and the systematic misclassification of fear, surprise and disgust as happiness. https://www.mdpi.com/2079-9292/15/4/849
  6. researchBond & DePaulo (2006), Personality and Social Psychology Review 10(3):214–234 — 206 documents, 24,483 judges; 54% overall accuracy, the 47%/61% split, and higher accuracy on audible than visible lies. https://journals.sagepub.com/doi/10.1207/s15327957pspr1003_2
  7. researchDePaulo et al. (2003), Psychological Bulletin 129(1):74–118 — 158 cues across 120 samples; many behaviours show no discernible or only weak links to deceit. https://psycnet.apa.org/record/2002-11509-006
  8. researchLuke (2019), Perspectives on Psychological Science 14(4) — simulation evidence that deception-cue effect sizes are inflated by publication bias and low power, and that real effects cannot be distinguished from false positives. https://journals.sagepub.com/doi/10.1177/1745691619838258
  9. researchConstâncio et al. (2023), PLOS ONE 18(2):e0281323 — systematic review of 81 machine-learning deception studies; 70% used mock deception, the largest real-life dataset holds 121 videos, and industry application is premature. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0281323
  10. researchKahneman & Beatty (1966), Science 154(3756):1583–1585 — the founding demonstration that pupil diameter tracks memory load. https://doi.org/10.1126/science.154.3756.1583
  11. researchMathôt (2018), Journal of Cognition 1(1):16 — pupillometry review — that anything increasing processing load dilates the pupil, that linking specific cognitive processes to the response is doomed to fail, and the factor-of-16 light response against 1–5% cognitive effects. https://journalofcognition.org/articles/10.5334/joc.18
  12. researchChen, Karabay, Mathôt, Bowman & Akyürek (2022) — concealed-information detection with research-grade pupillometry: the effect was not statistically significant for most participants individually. https://pubmed.ncbi.nlm.nih.gov/35867974/
  13. researchBarry & Wang (2023), Scientific Reports 13:13841 — iris melanin obscuring the pupil aperture on RGB cameras, the exclusion of dark-iris participants from prior work, and the 624% contrast improvement using far-red illumination. https://www.nature.com/articles/s41598-023-40796-0
  14. researchPRISMA-ScR scoping review, Frontiers in Digital Health (2026) — rPPG supports robust heart-rate estimation while HRV reliability degrades under motion, illumination variability and short windows; the 120-second and 60-second thresholds and the do-not-infer conditions. https://www.frontiersin.org/journals/digital-health/articles/10.3389/fdgth.2026.1855425/full
  15. researchBondarenko, Menon & Elgendi (2025), npj Digital Medicine 8:593 — demographic bias across 100 rPPG studies and the collated error rates by skin tone. https://www.nature.com/articles/s41746-025-01973-9
  16. officialISO/IEC 30107-3:2023 — presentation attack detection testing: the APCER and BPCER metrics, and the exclusion of attacks occurring other than at the capture device during presentation. https://www.iso.org/standard/79520.html
  17. industryiBeta published PAD testing methodology — Level 1 and Level 2 definitions: time per species, assumed attacker expertise, materials budget caps, permitted attack success and the 15% genuine-user rejection allowance. https://www.ibeta.com/iso-30107-3-presentation-attack-detection-confirmation-letters/
  18. officialNIST Face Technology Evaluations (FRTE / FATE) — the current evaluation programme, split into face-recognition and face-analysis tracks. https://www.nist.gov/programs-projects/face-technology-evaluations-frtefate
  19. officialNISTIR 8280 — FRVT Part 3: Demographic Effects (2019) — 189 algorithms from 99 developers over 18.27 million images; false-positive variation by factors of 10 to beyond 100; the sex, age and region findings; and the China-developed-algorithm result showing the differential tracks training data. https://nvlpubs.nist.gov/nistpubs/ir/2019/NIST.IR.8280.pdf
  20. officialNIST IR 8491 — FATE Part 10: Passive Software-Based PAD (2023) — 82 algorithms from 45 developers across nine attack categories, and the explicit exclusion of digital injection attacks from scope. https://nvlpubs.nist.gov/nistpubs/ir/2023/NIST.IR.8491.pdf
  21. officialRegulation (EU) 2024/1689 (EU AI Act) — the Article 3(39) definition including intentions, the Article 5(1)(f) prohibition, Recital 44, Annex III point 1 and point 4(a), the Article 50(3) and 50(5) transparency duties, and the Article 99 penalty tiers. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A02024R1689-20260727
  22. officialRegulation (EU) 2026/1744 (Digital Omnibus on AI) — the July 2026 amendment deferring entry into application of the Annex III high-risk obligations to 2 December 2027 and Annex I to 2 August 2028. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A32026R1744
  23. officialEuropean Commission — Guidelines on prohibited AI practices (February 2025) — the treatment of emotion recognition systems falling outside the Article 5(1)(f) prohibition, and the scope of “workplace”. This site could not open the document directly and does not quote it. https://ai-act-service-desk.ec.europa.eu/sites/default/files/2025-08/guidelines_on_prohibited_artificial_intelligence_practices_established_by_regulation_eu_20241689_ai_act_english_ied3r5nwo50xggpcfmwckm3nuc_112367-1.PDF
  24. officialRBI Master Direction on Know Your Customer — the V-CIP requirement for face liveness / spoof detection and face matching, and the permissive sentence on AI being used to make V-CIP robust. No provision on emotion, affect or deception appears in the Direction. https://www.rbi.org.in/Scripts/BS_ViewMasDirections.aspx?id=11566
  25. officialDigital Personal Data Protection Act, 2023 — the absence of a sensitive-personal-data or biometric category, the section 6(1) consent standard limited to what is necessary for the specified purpose, and the repeal of section 43A of the IT Act. https://www.meity.gov.in/static/uploads/2024/06/2bf1f0e9f04e6fb4f8fef35e82c42aa5.pdf
  26. officialDPDP Rules, 2025 — official release — the 18-month phased compliance timeline, the separate consent-notice requirement, the ₹250 crore maximum penalty, and the impact-assessment and stricter-checks duties on Significant Data Fiduciaries using new or sensitive technologies. https://www.pib.gov.in/PressReleasePage.aspx?PRID=2190655®=48&lang=2

Checked September 2026. The EU AI Act timeline was amended in July 2026; the date is part of the claim.

Ask an AI about this page

Opens your assistant with this page as the source, and a question rather than a summary. It will ask what you are building before it answers.

Nothing is sent from here. The link carries only this page’s title and address.