Two of the three problems on the laboratory's open list are, underneath, measurement problems. Telling a person from an automated client is one. Proving that a defence did not quietly damage the thing it defends is the other, and nobody publishes that number because nobody measures it. This role is the modelling and evaluation work behind both.
What you'll work on
Classification under conditions that move. The population on the other side is not static: it iterates, probes and adapts, so a model that is excellent against last month's distribution is a model whose evaluation has expired. You will build things that degrade visibly rather than silently, and that flag what they have not seen before instead of guessing confidently.
The harder half is the harness, not the model. The standard here is that a defence ships with its own regression on utility, published next to the number about attacks, which means you have to be able to measure the harm a filter does to legitimate work. That measurement does not exist as a solved practice anywhere; building it credibly is most of the contribution.
You will also work inside a constraint that shapes the whole pipeline. Business, corporate and API accounts are excluded from training by construction, not by a flag; special-category data is excluded entirely; and any personal account holder can switch their contribution off at any time, with no reason asked and effect from the moment they do. Provenance for every corpus is documented because law now requires it to be. If the first thing you would reach for is more data, this will be an uncomfortable seat.
What we're looking for
Someone who has trained and shipped models rather than fine-tuned them, and who can explain why one failed and not merely that it did. You have dealt with noisy labels, class imbalance, and the gap between a validation metric and production behaviour. You read a paper to decide whether the technique is worth implementing, not to cite it.
Time spent thinking about how models break under pressure matters more here than fluency in any particular framework: evasion, poisoning, extraction, and the ways an evaluation can pass for the wrong reason. Go for the serving path is useful and learnable on the job.
Years of experience are not a filter. The quality of one thing you built and can defend in detail is.
How to apply
Email careers@neuraphic.com with the subject line "AI/ML Engineer." Send a resume, links to work that shows your reasoning (papers, repositories, writeups), and a short note on one of the open problems: what you would measure first, and how you would know you were fooling yourself.