Mid-Senior Data Scientist (Life Sciences)

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Summary: Join a team of digital and software technology experts providing R&D and engineering services, making a difference in a career full of opportunities. Highlights: 1. Work at the heart of cutting-edge life sciences projects 2. Design and govern biomedical data models for scientific accuracy 3. Leverage machine learning and Gen AI to extract insights from scientific data At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world’s most innovative companies unleash their potential. From autonomous cars to life\-saving robots, our digital and software technology experts think outside the box as they provide unique R\&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same. **YOUR ROLE** We are looking for a Mid\-Senior Data Scientist with a strong background in biomedical sciences and data engineering to join our growing team in Portugal. In this role, you will work at the heart of cutting\-edge life sciences projects, helping to design, build, and govern the data foundations that power scientific discovery and product development. You will act as a bridge between scientific, engineering, and product stakeholders \- translating complex biological knowledge into robust, scalable, and FAIR data solutions. In this role you will play a key role in: * Designing and governing biomedical data models: leading source\-to\-canonical mapping, ontology alignment, and schema governance including versioning, changelogs, and downstream impact assessments to ensure data integrity and scientific accuracy across complex biomedical domains * Building and maintaining data pipelines: developing robust, schema\-driven pipelines in Python and SQL, performing exploratory data analysis, and implementing validation frameworks that support high\-quality, reproducible scientific workflows * Driving knowledge graph development: applying hands\-on experience with RDF, OWL, SPARQL, and property graph modelling tools such as Neo4j and GraphDB to build and enrich knowledge graphs that connect biomedical entities across diverse data sources * Applying machine learning and Gen AI: leveraging applied ML experience and familiarity with Gen AI tools (including text generation APIs, chatbots, and enterprise search solutions) to extract insight and value from scientific data at scale * Championing FAIR data principles: designing and delivering FAIR data products, leading harmonisation efforts across multiple source systems, and ensuring persistent identifiers and provenance are embedded into every data product * Aligning stakeholders across disciplines: driving alignment between scientific, engineering, and product teams through clear communication, structured documentation, and a solutions\-focused mindset that keeps complex projects moving forward * Working with biomedical ontologies and controlled vocabularies: applying deep knowledge of resources such as Ensembl, UniProt, and Gene Ontology, including judgment on when and how to extend or map them to real\-world data challenges **YOUR PROFILE** * MSc or PhD in Bioinformatics, Biomedical Engineering, Molecular/Cell Biology, Neuroscience, Genetics, or related field * Excellent stakeholder management — able to drive alignment across scientific, engineering, and product stakeholders * Data modelling and harmonization experience across complex biomedical domains including source\-to\-canonical mapping, ontology alignment, persistent identifiers, and provenance. * Strong Python and SQL skills; comfortable building data pipelines and performing exploratory data analysis * Applied data science / ML experience relevant to knowledge graph or scientific data work * Strong communication skills with both technical and business stakeholders * Critical thinking, intellectual curiosity, and impact\-driven mindset * Ability to adapt and manage priorities in fast\-paced environments * Fluent in Portuguese and English * Nice\-to\-have: + Prior experience in a pharmaceutical or biotech organization + Experience with data catalogue, metadata registry, or schema registry tooling + Data engineering fundamentals: pipeline architecture, schema\-driven automation, validation frameworks + Hands\-on experience with ML frameworks and model lifecycle (build, deploy, monitor) + Hands\-on experience with Gen AI models and tools, such as text generation APIs, chatbots, and enterprise search solutions. + Track record of leading schema governance: versioning, changelogs, tagged releases, downstream impact assessment + Experience designing FAIR data products and leading data harmonisation efforts across multiple source systems + Knowledge graph experience: RDF, OWL, SPARQL, and property graph modelling (Neo4j/GraphDB) + Experience with LinkML or equivalent schema modelling frameworks (classes, slots, ranges, constraints, cardinality, ontology bindings) + Strong command of biomedical ontologies and controlled vocabularies (e.g. Ensembl, UniProt, Gene Ontology), including judgment on when/how to extend or map them **WHAT YOU'LL LOVE ABOUT WORKING HERE** * Join a multicultural and inclusive team environment. * Enjoy a supportive atmosphere promoting work\-life balance. * Engage in exciting national and international projects. * Hybrid work. * Your career growth is central to our mission. Our array of career growth programs and diverse professionals are crafted to support you in exploring a world of opportunities. * Training and certifications programs. * Health and life insurance. * Referral program with bonuses for talent recommendations. * Great office locations. **ABOUT CAPGEMINI** Capgemini is an AI\-powered global business and technology transformation partner, delivering tangible business value. We imagine the future of organizations and make it real with AI, technology, and people. With our strong heritage of nearly 60 years, we are a responsible and diverse group of 420,000 team members in more than 50 countries. We deliver end\-to\-end services and solutions with our deep industry expertise and strong partner ecosystem, leveraging our capabilities across strategy, technology, design, engineering and business operations. Apply now!

Posted by

João Santos

Indeed · HR

Location

Similar jobs

Mid-Senior Data Scientist (Life Sciences) by Indeed in 2026 | ok.com