yoinka

Master thesis: Question Set Embeddings

Ericsson

Stockholm,Stockholm,SwedenMid
Sign in to applyVerified 2h ago
Location
Stockholm,Stockholm,Sweden
Work model
On-Site
Level
Mid
Posted
5h ago

Skills

LLMMachine LearningNLPPython

About this role

Join our Team 
 About this opportunity    This research project explores whether representing document sections through a small set of generated, answerable questions can improve semantic retrieval in Retrieval-Augmented Generation (RAG) systems. The proposed approach will generate up to five questions per document chunk, combine them into one representation, and store a single embedding vector. Its performance will be compared with conventional raw-text embeddings and, where practical, separate embeddings for each question.  The project will assess retrieval quality, storage requirements, indexing cost, and latency, with the goal of identifying a scalable and efficient alternative for RAG applications.

What you will do

Design and implement a retrieval pipeline that segments documents, generates and filters questions using an LLM, and creates question-set embeddings linked to the original text Build controlled experiments comparing raw-text, single question-set, and multi-question embeddings Evaluate the approaches using metrics such as Recall@K, Precision@K, MRR, and nDCG, together with vector-store size, indexing cost, and retrieval latency You will also analyze failure cases, including unrelated topics within a chunk, redundant or incomplete questions, and queries containing exact terms, numbers, identifiers, or specialist terminology. The outcome will be practical recommendations for using question-based embeddings in scalable RAG systems   The skills you bring   Strong foundations in Natural Language Processing, Information Retrieval, Machine Learning, or related fields Practical experience with LLMs, text embeddings, vector databases, and Retrieval-Augmented Generation (RAG) Familiarity with Python and common NLP/ML frameworks Understanding of semantic search, dense retrieval, and evaluation of metrics such as Recall@K and MRR Experience designing controlled experiments and analyzing retrieval performance Familiarity with recent research on question-oriented retrieval, document expansion, or dense passage retrieval is desirable Why join Ericsson? At Ericsson, you´ll have an outstanding opportunity. The chance to use your skills and imagination to push the boundaries of what´s possible. To build solutions never seen before to some of the world’s toughest problems. You´ll be challenged, but you won’t be alone. You´ll be joining a team of diverse innovators, all driven to go beyond the status quo to craft what comes next.   What happens once you apply? Click Here to find all you need to know about what our typical hiring process looks like.Encouraging a diverse and inclusive organization is core to our values at Ericsson, that's why we champion it in everything we do. We truly believe that by collaborating with people with different experiences we drive innovation, which is essential for our future growth. We encourage people from all backgrounds to apply and realize their full potential as part of our Ericsson team. Ericsson is proud to be an Equal Opportunity Employer.  learn more.   Primary country and city: Sweden (SE) || Stockholm Req ID:  790725

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Master thesis: Question Set Embeddings at Ericsson, Stockholm,Stockholm,Sweden | Yoinka