New model version V4 Apollo released! Read more →
All careers

AI team

ML / Data Engineer (Master Thesis / Internship)

Zurich, Switzerland · Hybrid · Internship

Build an augmentator that can make any recording sound like it was captured anywhere

A roughly 6-month Master thesis or internship with one goal: take any voice recording and regenerate it as if it had been captured somewhere else, through a different microphone, in a different room, over a different phone line. Get that right and every recording we hold turns into thousands of training examples. You will work both sides of it, growing the real-world recordings in Acoustic Atlas and training the generative model that learns from them.

About us

Aurigin.ai is a Zurich-born startup on a mission to restore trust in digital communication by protecting high-stakes organizations from AI-generated voice fraud in real time. Our deepfake detection is used to stop account takeovers, executive impersonation, and fabricated recordings before they cause damage.

Acoustic Atlas is our crowdsourced platform for collecting diverse real-world acoustic recordings. Contributors help us capture how devices, rooms, and replay paths change audio, which is what lets us train detectors that hold up outside the lab.

We are a team of engineers, researchers, and builders with backgrounds across tech, consulting, and startups. We thrive on collaboration, embrace a "move fast to launch and iterate" mentality, and share a vision of a safer, more transparent world in the age of AI.

Why we need you

Detectors fail in the real world for a boring reason: real calls do not sound like training data. A cheap headset, a reverberant meeting room, a compressed VoIP connection, or a clip replayed from a phone speaker all change a voice, and a model that has never heard those conditions gets them wrong. Acoustic Atlas collects real examples of exactly this, but no collection effort covers every combination of device, room, and channel that exists. So we want to learn the transformation itself: give a model a recording and a target acoustic scenario, get back the same speech as it would have sounded there. Training conditions then come from a generator instead of from more recording sessions. How to get there is genuinely open, and that is the thesis.

What you will do

  • Build the augmentator: a generative model that takes source audio plus a target acoustic scenario and returns that voice as it would have sounded in that setup
  • Make the scenario space open-ended: represent acoustic conditions so the model can produce combinations it has never seen, rather than replaying a fixed list of effects
  • Improve how Acoustic Atlas collects: the capture flow, the metadata, and the quality control that decides whether a contribution is usable for training
  • Get data in at volume, by whatever channel works: crowdsourcing, paid recording sessions, open datasets, a public competition. Pick the ones that buy the most coverage and run them
  • Prove it out: measure whether detectors trained on generated audio actually hold up better on replay and real calls, and say so plainly if they do not

About you

  • Pursuing or recently finished an MSc (or equivalent) in CS, EE, ML, or a related field
  • Strong Python and hands-on ML, ideally PyTorch
  • Interest in audio, speech, or signal processing, from coursework or your own projects
  • Comfortable with data work: cleaning, pipelines, experimentation, and measurement
  • Comfortable with an open problem: the goal is clear, the route to it is not, and you will help decide it
  • Happy doing ambitious research that still has to end up in a real training stack, not only in a thesis
  • Fluent in English, available for around 6 months, and able to be in the Zurich office at least two days a week

Nice to have

  • Experience with generative models, audio ML, room impulse responses, or room acoustics
  • Familiarity with speech enhancement, dereverberation, or channel simulation, which is close to the inverse of what we are building
  • Prior work on datasets, crowdsourcing, annotation, or getting people to contribute recordings
  • Web or product curiosity (Atlas is a live contributor-facing site)

Benefits

  • Paid Master thesis or internship
  • One problem that is genuinely yours, from getting the recordings in to the model that learns from them
  • Your work goes into the training stack behind a detector that runs in production, not into a report that gets filed
  • Flexible hybrid setup: work where you are most productive, with at least two days a week in the Zurich office

Ready to apply?

If you want to use your skills for good and help make communications trustworthy, send your CV and a short note to ai-careers@aurigin.ai.

Apply via email
Aurigin
Aurigin

Contact us by filling out the form below. We'll get back to you as soon as possible.