Computational Chemistry

Configuration Interaction: Mixing Electronic Excitations to Recover Electron Correlation

The Hartree–Fock wavefunction for the ground state of the H₂ molecule dissociates to a nonsensical mixture that is half H⁻···H⁺, raising the molecule's energy at long range by hundreds of kJ/mol above the correct two-neutral-atom limit. Adding a single doubly-excited determinant — promoting both electrons from the bonding σ_g to the antibonding σ_u* orbital — and letting the two configurations mix variationally fixes the dissociation limit exactly. That mixing is configuration interaction: the workhorse idea that a molecule's true state is a coherent superposition of many electronic configurations, not a single one.

  • Core ideaΨ = Σᵢ cᵢ Φᵢ — linear expansion in Slater determinants
  • OriginE. A. Hylleraas (1929, He); Boys/Shavitt formalized MO-based CI (1950s–60s)
  • Solved byDiagonalizing the CI Hamiltonian matrix Hᵢⱼ = ⟨Φᵢ|Ĥ|Φⱼ⟩ (variational)
  • Correlation energyE_corr = E_exact(nonrel.) − E_HF; ≈ 1 eV/electron pair (rough rule of thumb), ~−0.04 Eₕ for He
  • Common truncationsCIS (excited states), CISD (~90–95% of E_corr), CISDT, CISDTQ
  • Exact within basisFull CI (FCI) — factorial cost; ~10⁹ dets is a large FCI
  • Key defectTruncated CI is not size-consistent (Davidson correction, CEPA, CC fix this)

Interactive visualization

Press play, or step through manually. The visualization is yours to drive — try it before reading on.

Open visualization fullscreen ↗

Watch the 60-second explainer

A condensed visual walkthrough — narrated, captioned, under a minute.

What Configuration Interaction Is: A Basis of Determinants

Configuration interaction (CI) is the most conceptually direct way to go beyond the mean-field Hartree–Fock picture and recover electron correlation — the energy lowering that comes from electrons instantaneously avoiding one another, which a single-determinant, orbital-averaged potential cannot capture. The idea is to write the exact many-electron wavefunction as a linear combination of Slater determinants: Ψ = c₀Φ₀ + Σₐʳ cₐʳ Φₐʳ + Σ cₐᵦʳˢ Φₐᵦʳˢ + ⋯, where Φ₀ is the HF reference and the other Φᵢ are determinants generated by promoting electrons out of occupied spin-orbitals (labelled a, b, …) into virtual spin-orbitals (r, s, …).

The expansion is organized by excitation level relative to the reference: singles (one electron moved), doubles (two), triples, quadruples, and so on. Because a molecular calculation uses a finite one-electron basis set that spans, say, M spatial orbitals for N electrons, the number of possible determinants is finite but combinatorially huge. If every determinant is included, the method is called Full CI (FCI), and within that basis its energy is the exact nonrelativistic Born–Oppenheimer answer. FCI is the gold standard against which every approximate correlated method is benchmarked.

The key point is that CI is a linear variational method. The coefficients cᵢ are the unknowns, and they are found not by iterating orbitals (as in HF) but by diagonalizing a matrix. This makes CI transparent — you can read off exactly which configurations dominate a state — but also, in its truncated forms, subject to a specific pathology (size-inconsistency) that we return to below.

The Machinery: Diagonalizing the CI Hamiltonian

Applying the linear variational principle to Ψ = Σᵢ cᵢΦᵢ turns the Schrödinger equation into a matrix eigenvalue problem: Hc = Ec, where Hᵢⱼ = ⟨Φᵢ|Ĥ|Φⱼ⟩ are matrix elements of the electronic Hamiltonian between determinants. Diagonalizing H gives eigenvalues E that are variational upper bounds to the true energies — the lowest eigenvalue is the ground state, and higher eigenvalues approximate excited states, which is one reason CI is a natural excited-state method.

Two structural theorems make this tractable. First, the Slater–Condon rules tell you that ⟨Φᵢ|Ĥ|Φⱼ⟩ vanishes whenever the two determinants differ by more than two spin-orbitals, because the Hamiltonian contains only one- and two-electron operators. So the CI matrix is extremely sparse. Second, Brillouin's theorem (Léon Brillouin, 1934) states that matrix elements between the HF reference and any singly-excited determinant vanish, ⟨Φ₀|Ĥ|Φₐʳ⟩ = 0. The immediate consequence: singles do not couple directly to the ground state and contribute to the correlation energy only indirectly, through their coupling to doubles. This is why CID (doubles only) already captures most of the correlation, and why doubles are the dominant correction.

Because the number of determinants (∼10⁶ to 10¹⁰ or more) dwarfs anything that could be stored as a dense matrix, real CI codes never form H explicitly. Instead they use direct CI — Björn Roos's 1972 breakthrough — which computes the matrix–vector product Hc on the fly inside an iterative eigensolver, typically the Davidson algorithm (Ernest Davidson, 1975), which converges the few lowest roots of a giant sparse matrix without ever storing it. This algorithmic pairing, together with the graphical unitary group approach (GUGA) for spin-adaptation, is what made large-scale CI practical.

Truncation, Cost, and the CIS/CISD Hierarchy

Full CI scales combinatorially (factorial-like) with system size: the number of determinants for N electrons in M orbitals goes roughly as (M choose N/2)² for a closed shell. For a modest molecule in a decent basis this is astronomically large — FCI benchmarks are limited to a handful of electrons in a few dozen orbitals (state-of-the-art FCI has reached ∼10¹² determinants only with heroic parallel effort). So in practice CI is truncated at a fixed excitation level.

  • CIS (CI singles): includes only single excitations. By Brillouin's theorem it gives zero correlation for the ground state, so it is not a ground-state correlation method — but it is a cheap, size-consistent, O(N⁴)–O(N⁵) method for excited states, essentially the Tamm–Dancoff approximation to TD-HF, and a common first look at UV–vis transitions.
  • CISD (singles + doubles): the standard truncation, scaling O(N⁶). It recovers ∼90–95% of the correlation energy for small molecules near equilibrium.
  • CISDT, CISDTQ: adding triples (O(N⁸)) and quadruples (O(N¹⁰)) climbs toward FCI at steeply rising cost.

The tradeoff is stark. Each extra excitation level adds two factors of N to the scaling and pulls in a rapidly diminishing sliver of correlation energy. That poor cost/benefit ratio for truncated CI is precisely why the field largely shifted toward coupled cluster theory (Čížek, 1966), which recovers the same physics far more compactly via an exponential ansatz, and toward perturbative methods like MP2. CI's enduring niches are excited states, multireference problems, and as the exact FCI yardstick.

Worked Example: H₂ Dissociation and the Two-Configuration Fix

The cleanest illustration of why CI matters is the dissociation of H₂. In a minimal basis, HF gives one occupied bonding MO, σ_g = (1s_A + 1s_B)/√(2+2S), and one virtual antibonding MO, σ_u* = (1s_A − 1s_B)/√(2−2S). The restricted HF ground state is the single determinant |σ_g ᾱ σ_g β̄|. Expand it into atomic orbitals and you find it contains equal weights of the covalent terms (H·−·H) and the ionic terms (H⁺ H⁻ + H⁻ H⁺). At the equilibrium bond length (r_e ≈ 0.741 Å) this is a decent description, but as r → ∞ the ionic terms should vanish entirely — two neutral H atoms cost far less energy than an ion pair. RHF cannot remove them, so the RHF curve dissociates ∼0.29 Eₕ (∼180 kcal/mol) too high.

Now build a two-determinant CI: Ψ = c₀|σ_g σ_g| + c₂|σ_u* σ_u*|, mixing in the doubly-excited configuration where both electrons occupy the antibonding orbital. Diagonalizing this 2×2 CI matrix lets the ionic and covalent contributions decouple. At r_e the ground root is dominated by c₀ (c₂ small and negative); as the bond stretches, |c₂| grows until at dissociation c₀ = −c₂ = 1/√2. In that limit the wavefunction collapses to the pure covalent, two-neutral-atom form (1s_A·1s_B) with the ionic terms exactly cancelled — the correct products. This tiny CI recovers the static (nondynamic) correlation that a single determinant fundamentally cannot, and it is the prototype for all multireference chemistry: whenever a bond breaks, a diradical forms, or two configurations become near-degenerate, one reference determinant is not enough.

The same 2×2 logic underlies why species like ozone, the oxygen molecule's excited states, transition-metal spin states, and Woodward–Hoffmann-forbidden transition states demand a multiconfigurational treatment — their zeroth-order wavefunction is an irreducible mixture of configurations, and the mixing coefficients carry real chemical information.

The Size-Consistency Problem and Its Corrections

Truncated CI has a notorious flaw first sharply diagnosed in the 1970s: it is not size-consistent (nor the closely related size-extensive). Consider two He atoms infinitely far apart. CISD on the supersystem cannot describe simultaneous double excitations on both atoms, because that is a quadruple excitation of the pair — outside the singles-and-doubles space. But CISD on a single He atom captures its full pair correlation. The result: E_CISD(He···He at ∞) > 2·E_CISD(He), a spurious error that grows with the number of electrons and makes truncated CI increasingly unreliable for large systems and for comparing molecules of different size.

The root cause is the linear ansatz. Coupled cluster's exponential form, exp(T̂)|Φ₀⟩, automatically generates the disconnected product excitations (½T̂₂² supplies exactly those simultaneous doubles) and is therefore rigorously size-extensive — the single biggest reason CCSD supplanted CISD as the correlated method of choice. Several fixes were engineered for CI:

  • Davidson correction (Langhoff and Davidson, 1974): the '+Q' estimate ΔE_Q ≈ (1 − c₀²)·E_corr(CISD) approximately restores the missing quadruples contribution; CISD+Q is a cheap, widely used patch.
  • CEPA (coupled electron-pair approximation; Meyer, Kutzelnigg, 1970s) and the averaged coupled-pair functional (ACPF) modify the CI equations to approximate size-extensivity directly.
  • MRCI+Q: multireference CI with a Davidson-type correction remains one of the most accurate practical methods for excited states and bond-breaking, at high cost.

These corrections are empirical or approximate — a reminder that truncated CI's error is structural, not just numerical, and that the community's move to coupled cluster and to multireference perturbation theory (CASPT2, NEVPT2) was driven by exactly this defect.

Multireference CI and Where CI Lives Today

Modern quantum chemistry rarely runs a single-reference CISD as a production method, but CI ideas are everywhere. The dominant framework for strongly correlated, near-degenerate systems is the Complete Active Space Self-Consistent Field (CASSCF) method of Roos, Taylor, and Siegbahn (1980), which is a Full CI within a chosen active space of chemically important orbitals — for example CASSCF(6,6) means an FCI of 6 electrons in 6 orbitals — while the orbitals themselves are optimized. CASSCF captures static correlation (bond breaking, diradicals, metal d-manifolds); dynamic correlation is then layered on top by CASPT2 (Andersson, Malmqvist, Roos, 1990) or NEVPT2, or by MRCI built on the CAS reference. This CASSCF/CASPT2 pairing is the standard tool for excited-state photochemistry, conical intersections, and transition-metal spectroscopy.

Even the abandonment of literal full-CI has been partly reversed by algorithmic advances. Density Matrix Renormalization Group (DMRG) (White, 1992; adapted to quantum chemistry by White and Martin, 1999) can now treat active spaces of ∼50 orbitals — far beyond conventional CASSCF's ∼18-orbital ceiling — by compressing the CI vector as a matrix-product state. Selected CI methods (CIPSI, dating to Malrieu's 1973 work, and its modern revivals like Heat-Bath CI and adaptive sampling CI, ∼2016–2017) recover near-FCI accuracy by iteratively keeping only the determinants with the largest coefficients, exploiting the fact that most of the factorial FCI space is chemically dead weight.

The historical arc is worth noting: Egil Hylleraas's 1929 explicitly correlated treatment of helium and the earliest determinant-mixing work predate Hartree–Fock's dominance; Boys and, later, Shavitt built the MO-based CI apparatus in the 1950s–60s; Roos's direct CI (1972) and Davidson's diagonalizer (1975) made it scalable. CI is thus both the oldest correlated method and, through selected-CI and DMRG, one of the most actively developed — the enduring language in which we describe what it means for a molecule to be more than a single electronic configuration.

Truncated CI versus Coupled Cluster at the singles-and-doubles level
PropertyCISDCCSD
Wavefunction formLinear: (1 + Ĉ₁ + Ĉ₂)|Φ₀⟩Exponential: exp(T̂₁ + T̂₂)|Φ₀⟩
Variational (E ≥ E_exact)Yes — energy is an upper boundNo — not strictly variational
Size-consistentNo (error grows with system size)Yes
Includes disconnected quadruplesNoYes, via ½T̂₂² term
Formal scalingO(N⁶)O(N⁶)
% correlation recovered (small mol.)~90–95%~95–99%

Frequently asked questions

Why does CI singles (CIS) give no correlation energy for the ground state but is still useful?

By Brillouin's theorem, the matrix elements ⟨Φ₀|Ĥ|singly-excited⟩ all vanish, so single excitations do not mix with, or lower, the HF ground state — CIS returns exactly E_HF for the ground state. Its value is entirely for excited states: diagonalizing the singles block yields excitation energies (it is essentially the Tamm–Dancoff approximation to TD-HF), giving a cheap, qualitatively reasonable first picture of UV–vis transitions.

What is the difference between size-consistency and size-extensivity?

Size-consistency is the specific requirement that the energy of two non-interacting fragments A and B at infinite separation equals E(A) + E(B). Size-extensivity is the more general, mathematically precise condition that the correlation energy scales linearly with the number of electrons for a homogeneous system. Truncated CI (e.g., CISD) violates both; coupled cluster and MP perturbation theory satisfy them, which is the core reason those methods are preferred for anything but small molecules.

How is CISD different from CCSD if both use only singles and doubles?

The excitation operators are the same, but the ansatz differs. CISD is linear, (1 + Ĉ₁ + Ĉ₂)|Φ₀⟩, whereas CCSD is exponential, exp(T̂₁ + T̂₂)|Φ₀⟩. Expanding the exponential generates disconnected higher excitations — the ½T̂₂² term produces the simultaneous double-doubles (a quadruple) that CISD entirely lacks. That single structural difference makes CCSD size-extensive and typically more accurate, at essentially the same O(N⁶) formal cost.

When do you actually need a multireference CI instead of single-reference methods?

Whenever a single Slater determinant is a poor zeroth-order description — i.e., when two or more configurations are near-degenerate. Classic triggers are homolytic bond dissociation, diradicals and diradicaloids (ozone, TMM), open-shell transition-metal complexes with close-lying spin states, and points near conical intersections in excited-state chemistry. A diagnostic is the CASSCF weight of the leading configuration or the CCSD T₁ diagnostic (>~0.02 signals multireference character).

Why is Full CI called the 'exact' answer if it still has errors versus experiment?

Full CI is exact only within its two limitations: the finite one-electron basis set and the nonrelativistic Born–Oppenheimer Hamiltonian. It exactly diagonalizes the electronic Hamiltonian in the given basis, so it recovers 100% of the basis-set correlation energy. Remaining discrepancies with experiment come from basis-set incompleteness (addressed by extrapolation to the complete-basis-set limit), neglected relativistic effects, and non-Born–Oppenheimer/vibronic coupling — not from the CI treatment itself.

Does the Davidson correction ever fail or overcorrect?

Yes. The +Q correction, ΔE_Q ≈ (1 − c₀²)·E_corr, is an estimate of the missing disconnected quadruples that assumes the reference weight c₀² is meaningful and near 1. It becomes unreliable when the wavefunction has strong multireference character (c₀² well below ~0.9), where it can substantially overshoot, and it is not variational — CISD+Q energies can dip below the true energy. For strongly correlated systems one uses rigorously size-extensive methods (CC, CEPA/ACPF, or MRCI with a size-extensivity-corrected functional) rather than trusting +Q.