Master's Thesis: What Should an AI Agent Log for Failure Attribution

Ericsson•Published 5 hours ago•First seen 5 hours ago

Join our Team

About this opportunity:

The operation of mobile networks is becoming increasingly autonomous and complex. AI agents are already being introduced into network operations to collect evidence, form diagnostic hypotheses, and recommend corrective actions.

These systems are typically organised as pipelines in which each agent has a defined task. For example, one agent gathers evidence, another reasons about the fault, and a subsequent agent proposes a diagnosis, logs, or a corrective action. When the final output is wrong or uncertain, an important question arises: was the failure caused by the uncertainty of the first agent, an independent failure in the second agent, or an interface between them that lost information?

The fundamental question of this thesis is what the agent and its surrounding platform must log so that a failure can be explained after the event without excessive cost, complexity, or exposure of sensitive network data.

The objective is to identify and experimentally evaluate the minimum logging information required to attribute failures in a selected telecommunications use case. The work corresponds to one or two students, 30 hp each. The location is Stockholm/Kista, and the preferred starting date is early January.

What you will do:

  • Conduct a focused literature study on observability in LLM-based multi-agent systems, failure attribution, and error taxonomies.
  • Perform a logging-ablation experiment by systematically removing log categories and measuring the effect on root-cause and first-error attribution.
  • Analyse the trade-off between attribution value and logging cost, including trace size, latency, privacy, and operational usefulness.
  • Propose a minimum sufficient logging profile for telecom agents.
  • Validate the proposed profile on a small holdout set or through review with telecommunications subject-matter experts.
  • Document the results, limitations, and recommendations for future development.

The skills you bring:

  • You are studying Computer Science, Data Science, Computer Engineering, Electrical Engineering, or a related field.
  • You have solid Python programming skills.
  • You are interested in careful experimental design, observability, and AI systems.
  • You have good analytical, problem-solving, and technical writing skills.

The following knowledge or experience is considered a plus:

  • Large language models or agent frameworks.
  • Network operations or fault management.
  • Software testing and failure analysis.
  • AI governance, privacy, or responsible AI.
  • Statistics, structured logging, or autonomous networks.

Keywords: machine learning, statistics, Python, AI agents, structured logging, failure attribution, and autonomous networks.

Why join Ericsson?At Ericsson, you´ll have an outstanding opportunity. The chance to use your skills and imagination to push the boundaries of what´s possible. To build solutions never seen before to some of the world’s toughest problems. You´ll be challenged, but you won’t be alone. You´ll be joining a team of diverse innovators, all driven to go beyond the status quo to craft what comes next.
 
What happens once you apply?Click Here to find all you need to know about what our typical hiring process looks like.Encouraging a diverse and inclusive organization is core to our values at Ericsson, that's why we champion it in everything we do. We truly believe that by collaborating with people with different experiences we drive innovation, which is essential for our future growth. We encourage people from all backgrounds to apply and realize their full potential as part of our Ericsson team. Ericsson is proud to be an Equal Opportunity Employer. learn more.

Primary country and city: Sweden (SE) || Stockholm

Req ID: 791201