New Jersey has told educators that its new statewide assessment is adaptive, precise, and carefully engineered. But as teachers begin working with the system, a basic question keeps coming up: how does the test decide which question a student sees next?

For teachers, understanding how an assessment works is part of understanding what its results mean. If two students in the same classroom answer different questions, and those questions feed into scores used to describe academic performance, educators have a reasonable interest in knowing how those decisions are made, how the questions are calibrated, and how the system ensures fairness.

New Jersey's NJSLA-Adaptive (NJSLA-A) has provided some answers, but it has also exposed a broader challenge: how much of the machinery behind an algorithmic assessment should the people who use its results be able to see?

The engine is not just easier or harder questions

It's tempting to describe adaptive testing as a simple seesaw: answer correctly and the next question gets harder, answer wrong and it gets easier. That's not entirely wrong, but it misses the statistical engine running underneath.

Psychometricians don't simply look at the percentage of students who get a question right. They rely on a framework called Item Response Theory (IRT), under which every question carries three key measurements:

  • Difficulty — how hard is this item, statistically speaking?
  • Discrimination — how well does this item separate a struggling student from a strong one?
  • Guessing — what is the probability that a low-ability student could get it right by luck?

The NJSLA-A's algorithm uses these pre-calculated measurements to estimate a student's ability in real time. The New Jersey Department of Education's NJSLA-Adaptive FAQ says the assessment begins with items of medium difficulty and adjusts the difficulty of subsequent groups of questions based on how students perform, not after every single click, but once enough evidence has accumulated to refine the estimate.

The algorithm doesn't invent the assessment universe from scratch, either. It operates within a professionally constructed item pool that has already been vetted by humans.

A three-dimensional illustration on a violet background: a black graduation cap resting on a stack of three books, with a magnifying glass, a light bulb and an alarm clock arranged around it.

Where the questions actually come from

The path from a rough draft to a live operational question is long. Rather than listing every procedural step, it's easier to think of the process as four broad gateways:

  • Development and content alignment — writers and NJDOE content experts draft questions explicitly tied to the New Jersey Student Learning Standards.
  • Bias and sensitivity review — committees scrutinize every item for cultural, gender, and socioeconomic bias to ensure fairness across student populations.
  • Field testing and statistical analysis — items are piloted with students to gather real performance data. This is where the IRT measurements are calculated.
  • Final calibration and operational deployment — items that perform well statistically are approved for use in the operational adaptive engine.

New Jersey educators participate at multiple points along this continuum, reviewing passages, validating rubrics, and evaluating alignment. The algorithm doesn't write the questions; it selects them from a carefully curated bank.

The field test was necessary, and it was not free

The Fall 2025 field test is a case study in the tension between psychometric necessity and classroom reality.

For psychometricians, field testing is standard practice — you cannot deploy thousands of new items without testing them first. NJDOE told districts plainly, in its August 2025 field-test broadcast, that field-test results would not be provided to students, educators or schools because the purpose was “to evaluate the quality and performance of test items—not to measure student achievement”; it was gathering data to evaluate item validity, reliability, and fairness.

For the teacher who administered that field test, though, it represented a real cost. They spent class time on an assessment the department had already said would yield no feedback for them, their students, or their school. Teachers were asked to help students get comfortable with a new platform while being told the data generated wouldn't be reported back to them.

Transparency here means the state openly acknowledging that cost, rather than simply pointing to statistical necessity. Teachers deserve a clear explanation of how the field-test data was used to refine the algorithm, and whether any items were discarded or revised based on what it revealed.

Overhead view on a bright blue background of crayons, coloured pencils and a ruler scattered outwards from a miniature yellow school bus at the centre.

How a raw ability estimate becomes a Level 3

This next part is where the transparency conversation gets murkier.

The adaptive algorithm produces a precise mathematical estimate of a student's ability, often called a theta score. But that raw number is meaningless to parents and educators until it's translated into the familiar proficiency labels: Level 1, 2, 3, 4, or 5.

That translation doesn't happen automatically. It happens through a separate process called standard-setting, where panels of educators meet after the test is administered to decide where the lines are drawn on the ability scale — where "Proficient" begins, where "Advanced" starts.

How those panels are selected, what data they're shown, and how they map the algorithm's numbers to a performance level remains important, and largely opaque to the public. If the adaptive test changes the underlying scale but the cut scores are simply adjusted to maintain the status quo, parents and school boards may never know whether performance genuinely improved or whether the bar was simply moved.

What is actually known about AI scoring

The transparency debate becomes even more charged when it comes to writing, because here, what began as teacher speculation has been confirmed.

In early 2026, teachers in public forums were asking whether NJSLA writing responses would be scored by artificial intelligence or by human readers. The answer, according to public reporting by Government Technology and other outlets, is both: New Jersey is using automated scoring for the written portions of the NJSLA and NJGPA beginning with the Spring 2026 administration. The scoring engine was reportedly trained on student responses from fall practice administrations; most responses are scored automatically, while ones flagged as unusual or borderline are routed to trained human scorers. Cambium, the state's assessment vendor, has estimated that roughly a quarter of responses would be routed for hand-scoring.

That figure is worth holding against NJDOE’s own explanation for why results are late. The department’s July 2026 standard-setting presentation attributes the delay to “the extensive hand scoring of the ELA writing prompts this year, and the standard setting requirements.” A contract assumption of roughly one response in four, and a public explanation resting on extensive hand scoring, are not obviously the same picture. There may be a plain reconciliation — a first-year ramp, a wider review threshold, a definition that counts flagged-for-review volume — but none has been published, and the question a district can fairly ask is what share of its students’ writing was read by a person.

The assessment itself is straightforward: NJDOE's FAQ says the NJSLA-A writing component consists of a single extended-response task scored with a holistic rubric on two dimensions, Composition and Conventions, and the department has published the official rubrics.

There's also a cautionary tale worth keeping in mind. Massachusetts discovered that roughly 1,400 essays had been incorrectly scored by an automated system — a different vendor's system than the one New Jersey uses, but exactly the kind of failure that human-review safety valves exist to catch, and part of the reason flagging procedures matter so much.

So the useful question is no longer "Is AI scoring it?" It's how the automated system was validated, how human oversight works in practice, and what happens when the system produces an uncertain result.

Educators should understand the rubric, the training, and the quality-control procedures behind every score, automated or human. The same goes for rescoring. Teachers remember that schools previously had mechanisms for having certain responses rescored, and with automated scoring now confirmed, that mechanism matters more, not less. A rescoring pathway is more than an administrative detail — it's a safety valve, a signal that the system acknowledges its own fallibility.

Looking down a narrow aisle between two tall library shelves packed tight with books, a red carpet running away to a further shelf at the end.

Transparency does not mean publishing the source code

There's an important misconception on the other side of this debate, too.

Demanding transparency does not mean New Jersey should publish the source code for the adaptive engine. There are legitimate security reasons for protecting the operational item pool. If educators knew precisely how every item-selection rule worked, it could make the system easier to game, and releasing every operational question would compromise the assessment itself.

So the transparency standard shouldn't be show us everything. It should be enough for people to independently understand why they should trust the system, including clear technical documentation explaining:

  • How item difficulty, discrimination, and guessing are established under IRT
  • How items are field-tested and revised
  • How the adaptive pathway is constructed while maintaining content balance
  • How different adaptive pathways are placed on the same comparable scale
  • How the standard-setting panels are selected and how cut scores are determined
  • How scoring quality is monitored — automated, human, and the boundary between them
  • How unusual or questionable responses are handled
  • How items are removed or revised when they perform poorly

Educators don't need the source code to understand the methodology. But they do need enough information to evaluate it.

The real transparency test comes after the test

Perhaps the most important point is that transparency cannot end with a FAQ.

Spring 2026 was the first operational administration. The state now has something it did not have before: actual operational data.

NJDOE says the Spring 2026 results will be released in fall 2026, after standard-setting and approval of the new performance cut scores. That creates a real opportunity.

New Jersey can show educators not just what the assessment is supposed to do, but how well it actually did it — publishing technical documentation explaining how the new scale was established, how items performed, how the adaptive model performed across student groups, how the automated scoring performed against its human-review checks, and how the new scores relate to previous assessment results.

That would transform transparency from a communications exercise into a form of public accountability.

The teacher's question is reasonable

At the heart of this debate is a simple concern: teachers are being asked to prepare students for a system they cannot fully see.

That's a fair position. Teachers don't need to become psychometricians, and they don't need to understand every equation inside the scoring model. But they should understand enough to explain the assessment to students, interpret the results responsibly, and recognize the limits of what those results mean.

At the same time, skepticism shouldn't tip into misinformation. The fact that educators have questions doesn't prove the algorithm is flawed, and the fact that the state has provided technical explanations doesn't mean every concern has been resolved. Both can be true at once.

The standard New Jersey should aim for

The best standard for NJSLA-A transparency is neither blind trust nor total disclosure. It is verifiable trust.

New Jersey should be able to say:

  • Here is how the IRT engine works.
  • Here is how we field-tested the items, and here is what we learned.
  • Here is how we know different students' adaptive pathways produce comparable scores.
  • Here is how the standard-setting panels were chosen and how they determined the cut scores.
  • Here is how writing is scored — by machine and by human — and here is how we monitor quality.
  • Here is what the results mean, and here is what they do not mean.

And perhaps most importantly: here is the evidence.

That is the point at which an adaptive assessment stops being a black box, not because every line of code has been published, but because the people responsible for interpreting the results have enough information to understand and challenge the system intelligently.

New Jersey has built an impressive engine. The question that matters now isn't whether the algorithm can crunch the numbers, but whether the state is willing to open the hood far enough — the IRT mechanics, the field-test costs, the standard-setting logic, and the scoring safeguards — for the educators who actually teach and learn to trust what they see.


Related reading: The Great NJSLA-Adaptive Debate · From PARCC to Adaptive