Data Scientist/ RAG Specialist
Are you interested in working with the World’s leading AI-First Quality Engineering Company? Ready to advance your career, team up with global thought leaders across industries and make a difference every day? Join us at QualityAI!
We are looking for a Data Scientist/ RAG Specialist to join our growing team in Canada!
Role: Data Scientist/ RAG Specialist
Location: CAN/MEX/ARG(remote)
Must have Skills:
- Hands-on experience configuring and tuning hybrid BM25 and vector search in Elasticsearch, OpenSearch, or a comparable search platform.
- Strong knowledge of search engineering concepts, including indexing, retrieval, scoring, ranking, relevance, and production search system tradeoffs.
- Semantic BM25 and ANN search with Cross-encoder and specifically dealing with how to handle Recency and Relevancy in search.
- Strong communication skills are extremely important
Why this role matters:
The Search Data Scientist will design, evaluate, and improve search and discovery solutions that help users find relevant Client journalism quickly and reliably across large, diverse content collections. As a member of cross-functional project teams, the Search Data Scientist will combine data science, information retrieval, and search engineering expertise to tune production search systems and deliver measurable improvements in relevance, quality, and performance.
The team works closely with product, engineering, editorial, and other functions across the organization to define search quality, build robust evaluation practices, and develop scalable retrieval and ranking solutions.
This is for Text, Video and Image search with Meta Data.
What you will do:
- Configure, tune, and maintain hybrid search systems that combine BM25 or other lexical retrieval methods with vector and embedding-based retrieval.
- Design search experiments and evaluation frameworks, including offline relevance judgments, benchmark datasets, online testing, and analysis of quality and performance at scale.
- Develop and validate retrieval and ranking algorithms across large datasets, measuring improvements with appropriate relevance, coverage, latency, and business metrics.
- Build and tune query classification, intent detection, and routing approaches that direct queries to the most effective retrieval strategy, index, or downstream workflow.
- Develop post-retrieval reranking, filtering, deduplication, diversification, and elimination strategies that improve precision and result usefulness.
- Analyze search behavior, failure modes, and relevance gaps; synthesize findings into clear, actionable recommendations for product and engineering teams.
- Partner with search engineers and cross-functional teams to optimize indexing, retrieval, ranking, and serving approaches for accuracy, efficiency, scalability, and cost.
- Establish repeatable approaches for relevance tuning across large and evolving content collections, accounting for freshness, metadata quality, content type, and query intent.
- Support the full search development lifecycle, from problem definition and prototyping through integration, deployment, monitoring, and iteration.
- Communicate analysis and present findings clearly, adapting to a range of technical and business audiences.
- Implement and regression test iterative algorithm improvements based on user and stakeholder feedback.
Who you are:
- 3+ years of relevant data science, search, or information retrieval experience, with strong proficiency in Python, and experience working with large-scale semi-structured data.
- Bachelor’s degree in Data Science, Computer Science, Information Retrieval, or a related field.
- Strong knowledge of search engineering concepts, including indexing, retrieval, scoring, ranking, relevance, and production search system tradeoffs.
- Hands-on experience configuring and tuning hybrid BM25 and vector search in Elasticsearch, OpenSearch, or a comparable search platform.
- Experience with embedding-based retrieval, approximate nearest-neighbor search, vector index configuration, and lexical-vector score combination or fusion strategies.
- Experience developing and testing query classification approaches, including query classification, intent detection, query expansion, and routing.
- Experience with post-retrieval techniques such as learning-to-rank or model-based reranking, filtering, deduplication, thresholding, and result diversification.
- Demonstrated ability to validate search algorithms at scale using offline and online evaluation methods, including relevance judgments, benchmark sets, A/B testing, and statistical analysis.
- Familiar with ML engineering and ML Ops practices, with a track record of delivering runtime solutions
- An effective communicator who can tailor analysis and presentations to both technical and non-technical audiences
- Advanced-level professional competency in written and spoken English
What will set you apart:
- Experience in news media or working with news as data strongly preferred
- Master's degree in Data Science or a related field
- Eagerness to learn the technical nuances of large-scale media operations and identify opportunities within evolving systems
If you like what you have read, send us your resume and let’s start talking!
- Intrigued to find more about us?
- Visit our website at QualityAI (Formerly Qualitest Group) | World’s Leading AI-First Quality Engineering Company
- LinkedIn: https://www.linkedin.com/company/qualityaigroup
Nearest Major Market: San Jose
Nearest Secondary Market: Palo Alto