LINGUISTIC INFORMATION IN TRANSFORMER REPRESENTATIONS: A CRITICAL REVIEW OF PROBING METHODS, LAYER-WISE PATTERNS, AND EVIDENCE

Authors

  • Samra Hameed Author
  • Dr Saima Akhtar Chattha Author
  • Shazma Khalid Author
  • Mifra Hussain Author

Keywords:

transformers, linguistic representations, probing classifiers, mechanistic interpretability, syntax, morphology, discourse

Abstract

The new paradigm of language modeling with transformers has made the hidden states of these models with context-dependent behavior seem to contain rich linguistic information, but claims about what the hidden states contain are frequently not as well supported by evidence. This article is a structured narrative review of methods and results of linguistic feature extraction from Transformer representations. The 21 primary works that were published or posted from 2017 to June 2026 were selected for the core synthesis, after relevant and source quality assessments, peer-review, and direct support of the articles cited for the synthesis's claims. There are three valid conclusions that can be drawn. First, there is surface, lexical, morphological, syntactic, semantic and discourse information that is often accessible from the internal state, but is dependent on the specific architecture, training objective, language, task, and probe. Secondly, BERT-family encoders often show a more or less hierarchical progression from surface information in early layers, to syntactic information in middle layers and to semantic and discourse-level information in later layers, which may not be a common pipeline for all large language models. Third, the typical cross-lingual and discourse abstractions are more easily accessed in intermediate layers, while recent evidence from multi-model studies indicates that inflectional features could be linearly accessible over wide layer ranges. But the accuracy of the probe does not prove that the feature is used causally; other tasks such as control, minimum description length estimates, matched baselines, cross-model replication and tests of intervention are needed for more robust conclusions. The review does not use cross-dataset performance rankings, but uses an evidence ladder to identify decodability, selectivity, robustness and causal use.

Downloads

Published

2026-09-08