Data and AI
Data Scientist
Navi Mumbai • Full time • We hire for ability, not years
Turn our clients' data into decisions they can act on.
Before you apply
- We're direct about how we work. Read this section first.
- This is client work. You will not own one product for three years. You will work across manufacturing, logistics, retail, public safety and transport, and the problem will change under you more than once a year.
- We value people who know what they don't know, and say so. If you're honest about your gaps, we'll invest in closing them. If you hide weaknesses or oversell your experience, you won't last here.
- The unglamorous part is most of the job. Getting to a model is the quick bit. Deciding what to measure, finding out the labels are wrong, and proving a number is real takes the rest of the time. If that sounds like drudgery rather than the actual work, this is not the right fit.
- We move fast. Growth happens outside comfort zones. If this sounds harsh, we're probably not a fit. If it sounds like the work you have been waiting for, keep reading.
The role
Allerin ships AI into places where somebody measures whether it worked. Computer vision on a factory line, licence plates at a gate, analytics a client runs their week on. Every engagement has a number attached to it and a gate the release has to clear.
You own that number. Not the dashboard about it, the number itself. What it should be, whether it is real, whether it moved, and what to say when it did not. When a client asks for AI for quality control, the job is turning that into escaped defects per shift, a recall target, a false positive budget the line can live with, and an honest statement of what the model cannot see.
What you'll own
- Framing the problem: what the client asked for is rarely what they need. You sit in that conversation and come out with something measurable.
- The evaluation set: built from the client's own site and their own failure cases, not a public benchmark. This is the single most valuable artefact on most engagements and it is yours.
- Error analysis: looking at what actually broke, in detail, repeatedly. Accuracy hides a failure mode that happens five percent of the time. You find it, name it, and decide whether it matters.
- The thresholds: a false positive that stops a production line and a false negative that reaches a customer cost very different amounts, and only the client's economics tells you which way to lean. You work that out and write it down.
- Drift and the feedback loop: noticing the model got worse before the client tells you, and knowing which part of the data moved.
- Saying the number out loud: including when it is worse than last month, and including when the honest answer is that this should not ship.
What you must arrive with
We do not count years and we do not require a degree. We look for evidence.
- You can get to the data and interrogate it. SQL, Python, and the instinct to distrust a summary statistic until you have seen the distribution behind it.
- You have designed a measurement, not just reported one. You can explain why you chose a metric and what it failed to capture.
- You have found something wrong in your own analysis before someone else did, and you can describe how.
- You can explain a technical result to someone who will make a decision with it and who does not care how the model works.
- You write clearly. A finding nobody understood did not happen.
A GitHub account is required to apply. A posting that says it does not count years has to count something else, and this is it. We are not counting stars or green squares. We are looking for something you built and can be questioned on, so a fork you have not committed to or a tutorial followed to the end does not help you. A small, unfinished, honest project does. We look at every link, including whether a repository is a fork and how it was actually built.
Everything else is welcome and none of it is required. Kaggle, papers, write-ups, notebooks, demos, talks.
Nice to have
- Computer vision in production, particularly the parts that are not the model: lighting, camera placement, annotation quality, class definitions
- Evaluation of LLM or agentic systems, meaning eval sets and judges rather than prompt tweaking
- Experiment design and causal inference where a randomised test was not available
- Working with a client directly rather than through a product manager
What we offer
- A number that matters, on work a client measures, rather than an analysis nobody reads
- Range. Several industries and several problem shapes in a year
- A team that argues about evidence and does not mind being wrong in public
- Investment in your growth if you're honest about where you need it
How we hire
Apply through the form. Seven questions, and four of them are drawn from a larger bank, so no two candidates get the same set and there is no list to prepare against. We would rather read a handful of real answers than twenty self-ratings.
If we talk, one round is a review. We hand you a notebook that reaches a confident conclusion from a dataset, with something wrong in it, and ask you two questions. Would you show this to a client, and how do you know. Bring whatever tools you normally use. We are watching how you decide you are finished, which is the part of this job that got harder rather than easier.
A human reads every application, and answers that read as model-generated are tested in depth in that conversation, where they do not survive. Everything you write, and everything you link us to, will be discussed there.
Apply
Start with your email address. We send you a secure link, then read your resume so you do not retype what is already in it. After that there are 7 questions about this role.
By applying you agree to the handling of your data set out in our applicant privacy notice.