Verified current Job

VLM Research Engineer (m/f/d)

We’re looking for a Research Engineer to push the limits of vision-language models for real-world video understanding. You’ll work on applied, state-of-the-art multimodal models and turn them into production pipelines used by cust...

Job Full source details
Almetra Careers Berlin Source published Dec 18, 2025 Verified 18 hours ago
✓ 92% verification score · Source: Almetra Careers · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
VLM Research Engineer (m/f/d) at Almetra Careers
EmploymentFull Time
DepartmentEngineering

Overview

We’re looking for a Research Engineer to push the limits of vision-language models for real-world video understanding. You’ll work on applied, state-of-the-art multimodal models and turn them into production pipelines used by customers. Your role Design and adapt vision-language and video models for scene understanding, temporal reasoning and activity / action recognition Build and maintain large-scale training and evaluation pipelines on GPU clusters Curate and augment video-text and action datasets, including synthetic labels and retrieval-based augmentation Develop robust benchmarks for video QA, instruction following and temporal understanding, and use them to drive iterative model improvements Cut and refactor model architectures for efficiency and deployability (compression, pruning, distillation) Deliver production-ready inference pipelines to product and customer teams, working c

Full job description

Description

Full Job Description

We’re looking for a Research Engineer to push the limits of vision-language models for real-world video understanding. You’ll work on applied, state-of-the-art multimodal models and turn them into production pipelines used by customers.

Your role

  • Design and adapt vision-language and video models for scene understanding, temporal reasoning and activity / action recognition

  • Build and maintain large-scale training and evaluation pipelines on GPU clusters

  • Curate and augment video-text and action datasets, including synthetic labels and retrieval-based augmentation

  • Develop robust benchmarks for video QA, instruction following and temporal understanding, and use them to drive iterative model improvements

  • Cut and refactor model architectures for efficiency and deployability (compression, pruning, distillation)

  • Deliver production-ready inference pipelines to product and customer teams, working closely with CV, platform and robotics engineers

You bring

  • Completed PhD (or equivalent research track record) in computer vision, machine learning, robotics or a related field

  • Strong background in video-centric deep learning: scene understanding, temporal / activity / action recognition, or video generation

  • Experience training and adapting large vision or VLM models (e.g. InternVL, Qwen-VL, DeepSeek-VL, similar stacks)

  • Proven work with multi-GPU training (PyTorch, distributed, mixed precision) and large-scale datasets

  • Solid engineering habits: clean Python, reproducible experiments, reliable data and training pipelines

  • Track record of moving research into usable systems (demos, internal tools, or productised features) in fast-moving teams

Nice to have

  • Publications at top-tier venues (CVPR, ICCV, ECCV, NeurIPS, ICLR, etc.) on video, multimodal learning or scene understanding

  • Experience with 3D/4D scene representations, action generation or embodied / sense-plan-act style projects

  • Inference optimisation: quantisation, TensorRT, model distillation, or deployment on constrained hardware

  • Prior experience in a startup or applied research lab environment

What we offer

Employee Share Options Program for all permanent employees*

An increasing benefits list: currently includes Urban Sports club and quarterly team retreats.

Be on the forefront in defining what artificial intelligence means in manufacturing

Gain hands-on experience in working in an AI-first software company

Supportive and inclusive culture that values diversity and promotes the advancement of underrepresented groups within the company

Collaborate with a diverse (currently more than 10 nationalities) and talented team, working on cutting-edge projects with real-world impact

Network with professionals and leaders in the field, opening doors to potential future career opportunities

We have a very flat hierarchy, open 360° feedback, and flexible working hours

Ethics⚖: We are committed to developing ethical AI software

Don't meet all the requirements?

Almetra is committed to creating a workplace that is diverse, fair, and inclusive. We encourage candidates from all backgrounds, even if they do not meet every qualification, to submit their application. We firmly believe that having a team with diverse perspectives only strengthens our company and drives innovation. Our commitment also extends to providing an accessible environment for everyone, including those with disabilities. Please let us know if you require any accommodations during the application process or while working with us, and we will do our best to support you.

*Only full-time, permanent roles are eligible for stock options. Part-time roles, contract roles, work-student, internships and freelance roles are not eligible for stock options.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Discovered directly from the employer’s public Ashby Job Postings API. The complete public role content and compensation metadata were normalized into safe candidate-facing sections. Complete structured details were extracted from the public authoritative source while preserving the original application link.

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Almetra Careers ↗

Browse current Job and Scholarship listings from Almetra Careers →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books