Written for New Jersey district and school leaders, testing coordinators, and anyone who will put Spring 2026 results on a board slide next to prior years.
Sometime this fall, a New Jersey district will build a slide with ten years of NJSLA proficiency rates on it, and the last bar will come from a different test. The NJSLA-Adaptive replaced the fixed-form NJSLA operationally in Spring 2026, and NJDOE expects results in fall 2026, later than usual, because the state has to set performance standards first and obtain State Board approval.
That slide is where a measurement question turns operational. If the 2026 bar sits lower than 2025, a board member will ask whether achievement fell or the test changed. The honest answer depends on evidence that has not been published.
Note the question is not whether the new assessment works. Items can be sound, the platform stable, the scoring defensible, and the new test still not interchangeable with the one it replaced. Those are separate claims, and only the second one carries a trend line. (Whether adaptive testing measures better is a different argument, taken up in The Great NJSLA-Adaptive Debate. This piece is only about whether the numbers can be lined up.)
Why the comparison is harder than it looks
Consider a school that has used the same mechanical scale for a decade and buys a digital one. To check the new scale, you weigh things whose weight you already trust: a filing cabinet, a box of paper, a chair. If both scales read 25 pounds for the chair, that is strong evidence, and the reason is easy to overlook. The chair did not change while you switched scales.
Now weigh students instead of chairs, in January on the old scale and in May on the new one. A student reads 120 in January and 128 in May. Did the student gain eight pounds, or does the new scale read eight pounds high? Both readings fit either explanation, and a combination of the two.
That is the problem in miniature. Achievement is not a stable object. Students learn, forget, mature, change schools, and receive different instruction between two administrations. When a fall score and a spring score disagree, the difference contains real change in the student and any difference between the instruments, and the two numbers alone will not separate them.
The same logic applies when the levels agree. A student who comes out proficient on the old test and proficient on the new one looks like evidence of two instruments agreeing. It is also consistent with a student who improved on a test that ran slightly harder, or slipped on one that ran slightly easier. A matching label is not evidence of a matching scale.
This does not make the comparison useless. It means the analysis has to account for a moving target, which is ordinary work in educational measurement. But it is work, and it has to be done and shown.
The trap: proving what you already assumed
There is a specific way this goes wrong, and it is worth naming because it arrives disguised as evidence.
Suppose the only argument offered for comparability is that roughly the same share of students came out proficient under both tests. That is a circular reasoning problem. Similar rates are what you would expect if the two scales correspond — and also what you would expect if the cut was placed where it would produce a familiar-looking number. The observation fits both explanations, so it separates neither. The reasoning closes a loop instead of testing anything.
It gets worse if cut scores are chosen, even informally, to land near prior proficiency rates. Then the trend line is flat by construction, and a district reading stability off that chart is reading back its own starting assumption.
What breaks the loop is a linking design settled in advance rather than an argument assembled afterward. The distinction is not that such a design reaches outside the tests for the truth — nothing does that, and no procedure yields a student’s real achievement for the two scales to be scored against. The distinction is that a design fixed beforehand can come out badly. It can show the scales do not correspond. That is what separates a test of an assumption from a restatement of it, and without it you risk just confirming your starting assumption, not actually verifying it. What such a design produces is evidence about how two measurement systems relate, and the question is always whether that evidence is strong enough to carry the interpretation placed on it. The specific designs are the subject of a later section.
To be precise about New Jersey specifically: the method NJDOE describes is criterion-referenced rather than distribution-matching. Educator panels judge what a student at each level should know against written Performance Level Descriptors; they are not asked to reproduce last year’s percentages. That avoids the crudest version of the loop. What it does not do on its own is establish how the new scale relates to the old one. That is a different job entirely, and the distinction is worth being exact about.
The labels themselves already changed
A concrete instance of the gap between a label and its meaning sits in NJDOE's own materials.
NJDOE's School Performance Reports reference guide lists the fixed-form NJSLA levels in the past tense: “Level 1: Did not yet meet expectations, Level 2: Partially met expectations, Level 3: Approached expectations, Level 4: Met expectations, Level 5: Exceeded expectations.” The NJSLA-Adaptive FAQ lists the new levels as Did Not Yet Meet Expectations, Partially Meet Expectations, Approaching Expectations, Met Expectations, Exceeded Expectations.
Two of the five shifted from past tense to present participle. That is almost certainly a drafting choice, and reading policy into a verb form would be a mistake. But it makes the point exactly: the label is a name, not a measurement. The two systems do not even use identical words, and identical words would have settled nothing. What settles it is the cut score behind the label and the evidence connecting the two scales.
One local detail underpins all of this: in New Jersey, proficient means Level 4 or 5. The same reference guide is explicit that students “are considered proficient if they have met or exceeded expectations … which means they have a performance level of 4 or 5 on the NJSLA.” So the figure on the board slide is the share of students at Level 4 and above, and the entire trend line rides on one boundary: the cut between Level 3 and Level 4. Move that single line slightly and the headline rate moves with it.
What the Fall 2025 field test established
The Fall 2025 administration was a field test, not an operational assessment. Per NJDOE's September 16, 2025 broadcast, participation was mandatory across most secondary grades, students responded to items aligned to the prior grade level, and no results went back to students, families, educators, or schools. Its stated purpose was gathering evidence of the validity and reliability of the items.
A field test earns its keep. It shows whether items behave as intended, whether any perform differently across student groups, whether the adaptive engine selects sensibly, and whether items can be calibrated for operational use. Those answers are prerequisites for an operational administration.
Notice what the design cannot deliver. Students took prior-grade-level items and received no scores. That does not put an old-test result and a new-test result side by side for the same student on the same content. Describing Fall 2025 as the state having given both tests to the same students and demonstrated equivalence describes something that did not happen.
Two different pieces of work
Standard setting and linking get conflated constantly, and the difference decides whether a trend line is legitimate.
Standard setting decides where the cuts fall on the new scale, meaning what score a student needs to be called Level 4. NJDOE's July 2026 standard-setting presentation lays out the method in five steps: write Performance Level Descriptors, build Ordered Item Booklets from actual Spring 2026 student responses, convene panels of New Jersey educators, run structured rounds of judgment and discussion, and converge on recommended cut points.
Linking or equating asks the separate question of whether the new score scale can be expressed in terms of the old one at all. This is what makes a ten-year trend line real rather than decorative, and it takes a design built for the purpose. The common ones are a set of items appearing on both forms, a group of students sitting both assessments, or an external measure the new scale can be checked against.
Two of those are easy to describe wrongly. Anchor items sit inside both forms — that is what makes them anchors — and a common group is measured by both tests. Neither reaches outside the pair for an independent reading. They are mechanisms for connecting two measurement systems statistically, not sources of truth about students.
The common-group design also carries the cost this article opened with. If the same students sit both assessments months apart, the thing being measured has moved between the two readings, and the analysis has to model that movement rather than assume it away. The closer together the two sittings, the smaller that problem gets — which is part of why a common-item design, where both forms are measured in a single sitting, is attractive when it can be arranged.
They can come apart, and here they visibly do. The same presentation is blunt about the discontinuity: the transition “introduces a new scoring scale,” and “prior fixed-form cut scores from the NJGPA and NJSLA no longer apply—new performance boundaries must be established before test results can be reported.” A state can run a rigorous standard-setting process, as this one appears to be, and still have no published basis for comparing this year's proficiency rate to last year's. Standard setting establishes where the lines fall on the new scale. It does not say how that scale relates to the old one.
How far the public evidence goes
Published: the five performance levels, the rationale for the change, the standard-setting method and timeline, and the field-test design and participation requirements. All of it is linked above.
Not published: any technical report establishing that NJSLA-Adaptive scale scores are comparable to fixed-form NJSLA scale scores. The FAQ does not mention linking. Neither does the standard-setting presentation, which walks through the department's method step by step and describes no component that connects the new scale to the old one. That silence is not evidence the work is missing, since technical documentation routinely lags the administration it describes. It does mean a district cannot cite one today.
Not knowable from outside: Cambium's item pool, item-selection logic, and internal calibration work. None of that is public, and reasoning from the outside is no substitute for the documentation.
What to do with the fall 2026 slide
The practical consequence is narrow.
When Spring 2026 results arrive, a district can accurately report how many of its students landed in each performance level. Absent published linking evidence, it cannot assert that a change from 2025 to 2026 measures a change in student achievement.
So show the 2026 results and mark the break in the series rather than drawing a continuous line across it. A footnote stating that 2026 reflects a different assessment with newly established performance standards is not a hedge. It describes what the data supports, and it is far easier to say in advance than to explain after a board has read a two-point drop as a decline in instruction.
If NJDOE publishes linking documentation before the presentation, the footnote changes. That is the thing worth watching for between now and then: not another explanation of what adaptive testing is, but the technical report saying how the new scale relates to the old one.
Worth saying plainly: this is the harder of the two problems standing between a district and a multi-year chart. The other one is mechanical, and column names that drift between years will stop a comparison before the psychometrics ever get a chance to. That one you can fix yourself. This one you cannot.
The question was never whether New Jersey can validate an adaptive assessment. It can, and an assessment as long-established as the fixed-form NJSLA is a reasonable reference to validate against — that is ordinary practice, not a workaround. The question is narrower, and answerable: does the evidence available to the public show that the new scale corresponds to the old one closely enough to carry a ten-year trend line? Today it does not, because that evidence has not been published. The validation has to be designed and then shown, because achievement is not a chair that can be set on two scales and read twice.
