Ground Truth.
AI, checked against the source.

Learn · Intermediate

Multi-task learning: when shared skills help, and when tasks fight

Multi-task learning trains a model on several task objectives while sharing some of its parameters between them. Related tasks can teach the model reusable representations, reducing duplicated learning and sometimes improving generalization. Negative transfer occurs when that sharing makes a task worse, so the important design question is which information should be shared and which should remain task-specific.

Consider a system that identifies objects in pictures, estimates depth, and recognizes scene types. Edges, shapes, and textures can help all three jobs. Training separate models from scratch repeats much of that work. A shared visual encoder with separate output heads lets each task contribute evidence about common features, while the heads translate those features into different answers.

Rich Caruana helped establish the modern multi-task learning framework. Sebastian Ruder’s survey of multi-task learning in deep neural networks explains the main sharing strategies and why related objectives can help. Hard parameter sharing uses a common representation with task-specific heads. Softer sharing keeps more separate components and encourages selected representations or parameters to align.

The analogy is a workshop serving several products. All teams may benefit from the same reliable measuring tools and cutting equipment. A shared tool becomes harmful if one team continually recalibrates it for a different material without accounting for the others. Sharing is valuable because work is reused, but the workshop still needs a rule for resolving incompatible requirements.

Training usually combines losses from the tasks. Each loss measures error for one objective, and a weighted sum supplies the overall update. The weights are important: a large numerical loss can dominate even if its task is not more valuable. Dataset sizes and sampling rates matter too. Showing one task ten times as often changes the effective priority even if all written loss weights are equal.

Gradient descent turns each task’s loss into a direction for changing the shared parameters. When two tasks favor similar changes, one update can help both. When their directions conflict, an update that reduces one error can increase another. A negative inner product between gradients is one signal of local conflict, not a complete test of whether the tasks are semantically incompatible.

Tianhe Yu and colleagues’ Gradient Surgery for Multi-Task Learning proposes modifying conflicting gradients before combining them. Its method projects away a component that points against another task’s gradient. Imagine two colleagues pulling a shared cart in different directions and removing the part of one pull that directly opposes the other. The cart can move with less local conflict, but that does not prove the route is correct.

Other approaches change task sampling, loss weighting, network structure, or routing. Separate heads preserve different output spaces. Adapters can reserve some task-specific parameters. Mixtures of experts can route examples through different components, though routing itself needs training and does not automatically eliminate interference. More isolation protects specialization while sacrificing some sharing and increasing engineering or storage cost.

Negative transfer differs from ordinary underfitting. A small model may perform poorly because it cannot represent all objectives. Negative transfer asks whether joint training hurts a task relative to a meaningful separate-training baseline. The comparison must account for data, compute, and parameter budgets. A larger joint model beating a smaller specialist does not by itself show that sharing was the cause.

It also differs from catastrophic forgetting. Forgetting usually concerns an old skill damaged by later learning. Negative transfer can occur while both tasks are actively trained together. The two problems can share conflicting updates, but their evaluation setups differ: one asks about preservation across time, the other about the cost or benefit of joint learning.

Today’s dossier offers a useful boundary through Where Does Jagged Competence Come From?. Ioannis Tsiokos studies a small recurrent network with known lower-level computations and a downstream task. The results distinguish information that a probe can read from information reliably used or preserved. That paper is not a general multi-task benchmark, but it explains why a shared representation’s apparent knowledge does not guarantee dependable use by every objective.

A practical evaluation tracks each task separately. Record single-task baselines, shared-training results, and subgroup errors. Test both common and rare cases. If the overall average rises while one high-consequence task degrades, the average does not settle whether the design is acceptable. Ablation studies help isolate whether gains come from shared features, extra data, changed budgets, or task-specific components.

The honest caveat is that task relatedness is not a fixed property visible from a label. Two language tasks can conflict; a language and vision task can reinforce one another through shared concepts. Multi-task learning is a method for testing and exploiting shared structure. It succeeds when the objectives and architecture make reuse productive, while the evaluation remains sensitive to the tasks that pay the price.

Key papers
An Overview of Multi-Task Learning in Deep Neural Networks — Ruder (2017)
Gradient Surgery for Multi-Task Learning — Yu et al. (2020)

Key questions

What makes training multi-task rather than merely diverse?

Multi-task training combines identifiable task objectives that share some model parameters. A large mixed dataset can support that design, but diversity alone does not specify the objectives or sharing structure.

How is negative transfer different from catastrophic forgetting?

Negative transfer is harm caused by sharing training across tasks, including tasks learned simultaneously. Catastrophic forgetting usually describes loss of an earlier skill during later learning; the mechanisms can overlap.

Can gradient surgery guarantee that every task improves?

No: it modifies conflicting training directions under a particular optimization rule. It does not guarantee better generalization, appropriate task grouping, or improvement on every evaluation.
Cite this

APA

Ground Truth. (2026, October 7). Multi-task learning: when shared skills help, and when tasks fight. Ground Truth. https://groundtruth.day/learn/multi-task-learning-and-negative-transfer.html

BibTeX

@misc{groundtruth:multi-task-learning-and-negative-transfer,
  title  = {Multi-task learning: when shared skills help, and when tasks fight},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/learn/multi-task-learning-and-negative-transfer.html}
}

Topics: training · multi-task-learning · representation-learning · optimization