Skip to main content

How Standards Proficiency Is Calculated

The two algorithms Forefront uses to turn question scores into standards proficiency, side by side, and the situations where each one can surprise you.

Every colored square, ring and bar in Forefront comes from the same question: given everything we know about this student, how proficient are they on this standard? Forefront answers it in one of two ways, and which one your district uses is a setting rather than something you choose per report.

This article walks through both, using one student’s standards wheel as the running example, and is deliberately honest about the places where each answer can look wrong at first glance.

Question Average (the default)

Roll-up

A standard’s score comes from

every question aligned to it or to anything beneath it

its own questions, averaged with each sub-standard’s score

What gets equal weight

each question

each sub-standard

A broad, recent assessment

can move a whole domain on its own

only refreshes the sub-standards it actually assessed

Turned on

everywhere, by default

per district, by Forefront

If you just want to know which one you’re in, skip to Which one is my district using?.

What both algorithms do firstlink

Neither algorithm looks at raw points. Both start by turning every score into a proficiency band, and both then do the same two things with dates.

Questions carry the alignment, not assessments. A question is aligned to one or more standards — often several frameworks at once, so the same interview task feeds a Number Sense progression, a state standard and a Common Core standard. Only aligned questions say anything about a standard.

Each question is one equal vote. A question’s score is converted to a band using that question’s own rubric ranges, and from then on only the band matters. A 12-point task and a 1-point task count the same, and a student who tops the band counts the same as one who barely reached it.

Bands are numbers 0–4 while the math happens.

Band

Value

Well Below Basic

0

Below Basic

1

Basic

2

Proficient

3

Advanced

4

Averages are rounded up at the halfway point when they’re turned back into a color: 2.5 and above is Proficient, 1.5 up to 2.5 is Basic, and so on.

More recent evidence counts for more. When a standard has been assessed more than once, the values are combined with a weighted average rather than a plain one. The weight falls off with the gap between an assessment and the most recent evidence in view:

Gap behind the newest evidence

Relative weight

Same day

100%

1 week

51%

2 weeks

31%

1 month

14%

3 months

3%

6 months

1%

Two things about that table surprise people. The gaps are measured against the newest assessment in view, not against today, so a body of evidence doesn’t quietly decay over the summer — it re-weights when new evidence arrives. And the fall-off is steep: with a fall screener and a midyear screener just over five weeks apart, midyear carries about 91% of the answer and fall about 9%.

You never have to take that on faith. Click any node of a standards wheel and the body of evidence lists each assessment with the Weight it was given.

Question Average, the default algorithm

For one assessment, a standard’s value is the average of the bands of every question aligned to that standard or to anything beneath it. That value is rounded to a band and stored. Across assessments, those per-assessment bands are combined with the recency weighting above.

So a domain is, in effect, a straight average of all its questions, and the deeper structure of the standards tree doesn’t change the arithmetic — only which questions are counted.

A student's standards wheel with Place Value reading Below Basic

Place Value: Below Basic

The Number Sense progression for one second grader. The inner ring holds domains, the middle ring clusters, the outer ring individual standards — and grey means nothing has assessed that standard yet.

Where it can surprise you

A broad, recent assessment can move a whole domain by itself. Because a domain counts every question beneath it, a screener that touches twelve standards writes a fresh value for every ancestor of every one of them. That new value is the most recent evidence, so it carries almost all of the weight, and a domain that a class has been working on all year can be repainted by a single morning’s data. This is the single most common complaint about the default algorithm, and it’s the reason the roll-up mode exists.

A domain reports on whatever it has the most questions about. Question counts are rarely even. If a domain has six questions about one sub-skill and one question each about two others, three-quarters of the domain is a report on that first sub-skill. Nothing is broken — but a domain reading Basic can mean “fine at the two things we barely asked about, weak at the thing we asked about six times.”

Values are rounded twice. Each assessment’s value for a standard is rounded to a band before the recency average runs, so a student at 1.49 and a student at 0.51 both enter the time average as Below Basic. Two students with visibly different papers can come out identical.

A standard with no aligned question stays grey, and stays grey no matter how much evidence its parent or its siblings have. “Not Assessed” always means literally that: no question in view is aligned to that standard.

Roll-up, the alternative

Roll-up mode inverts the order of operations. Each standard is scored only from the questions aligned directly to it, using the same banding and the same recency weighting. Then those scores are averaged up the tree: a standard’s value is the average of its direct sub-standards’ values, together with its own score if it has questions of its own — which sit alongside the sub-standards rather than above them, for the reason in Standards aligned above the leaves.

Three consequences follow from that, and they’re the whole point of the mode:

  • Averaging is structural. Two branches count equally, however many questions or leaves sit under each.

  • Values stay continuous on the way up. Each standard’s own per-assessment value is still banded, exactly as above, but the averages climbing the hierarchy are not rounded at every level — only the standard you’re looking at is turned back into a color.

  • A sub-standard with no evidence is skipped, not counted as zero. It stays grey, and its parent reads on the sub-standards that do have evidence.

The same student's standards wheel in a roll-up district, with Place Value reading Basic

Place Value: Basic

The same student, the same scores, the same lens — read by the roll-up algorithm. Place Value and its Computation, Estimation and Problem Solving cluster both come out a band higher.

Where it can surprise you

A domain can read higher — or lower — than the questions under it “feel”. Giving each branch equal say is exactly what makes a domain robust against an uneven question count, and it also means the domain is no longer the average of the question list you’re looking at. The worked example below is one of these.

One question can outweigh five. A sub-standard assessed by a single question counts as much as a sibling branch assessed by five — that happens in the worked example below. It is the intent, since the two sub-skills matter equally, but it does mean one unlucky task can hold a domain down.

Filling a gap can move a parent a lot. Because empty branches are skipped, a sub-standard’s first ever score doesn’t just fill in one grey wedge; it becomes a full equal voice in its parent. Expect a domain to shift the first time a previously unassessed sub-skill gets assessed.

The Weight column can look like it contradicts the answer. Recency is applied per standard, before anything is averaged up, so an older assessment is only competing with newer evidence about the same standard. If a sub-standard was assessed in the fall and never since, the fall screener is 100% of that sub-standard’s score — and that score is then a full equal voice in its parent, no matter how far back it was. So a parent can read Basic while the panel above it shows a Below Basic assessment at 91% and a Basic one at 9%. Those percentages describe how the two assessments weigh against each other within one standard; they are not the recipe for a parent’s number.

A roll-up district's body of evidence where Place Value reads Basic while its newest assessment reads Below Basic

Not a bug: Place Value is Basic even though the midyear screener (91% of the weight) reads Below Basic. Ones, Tens, and Hundreds was only ever assessed in the fall, so within that sub-standard the fall screener carries everything — and the sub-standard carries a third of the domain.

A cluster’s own questions are a sibling of its children, not a summary of them — and they lend nothing to those children. This one is worth its own section — see Standards aligned above the leaves.

Assessments that compute their own standards don’t roll up. A few assessments (DIBELS-style ones, for example) come with a bespoke calculation that authors a value at every level it cares about, including cluster and domain levels. Those values are read exactly as the assessment authored them; rolling them up would average an authored domain against the children it was derived from. So in a roll-up district a DIBELS composite still behaves like a DIBELS composite.

Same student, same evidence, two answers

Here is the Place Value domain for the student in the screenshots above. Eight questions across the fall and midyear screeners; three sub-standards, one of which has three sub-standards of its own.

Standard

Questions

Assessed by

Score

Computation, Estimation and Problem Solving → Mental Math

2

Midyear

Basic (2.0)

Computation, Estimation and Problem Solving → Create and Interpret Representations

2

Midyear

Basic (2.0)

Computation, Estimation and Problem Solving → Solve Written Problems

2

Midyear

Below Basic (1.0)

Magnitude and Comparison

1

Midyear

Below Basic (1.0)

Ones, Tens, and Hundreds

2

Fall

Basic (2.0)

(The counts add to nine because one midyear task is aligned to two sub-standards at once — Mental Math and Solve Written Problems — so it is evidence for both.)

Question Average counts the questions. The midyear screener’s six Place Value questions banded as Basic, Below Basic, Below Basic, Basic, Well Below Basic and Below Basic, which averages to 1.17 — Below Basic. The fall screener’s two came out Basic and Basic, so fall’s value is Basic. Midyear is just over five weeks newer, which earns it 91% of the weight against fall’s 9%, and the domain comes out at 1.09 — Below Basic.

Roll-up works up the tree instead. Computation, Estimation and Problem Solving averages its three sub-standards: (2.0 + 2.0 + 1.0) ÷ 3 = 1.67. Place Value then averages its three sub-standards: (1.67 + 1.0 + 2.0) ÷ 3 = 1.56 — Basic.

Neither number is a bug. The default answer says most of what we asked this student about place value, they got wrong. The roll-up answer says of the three things place value is made of, they’re shaky on one and basic on two. They disagree because the questions aren’t spread evenly over the sub-skills — and that is the situation the two algorithms are built to read differently.

In a roll-up district the body of evidence shows you this directly: expanding an assessment groups its questions by the sub-standard they belong to, with each group’s score beside it, so a domain’s number can be traced to the groups it averaged rather than to a flat list of questions.

The body of evidence for Place Value, with the midyear screener's questions in a flat list

The default algorithm: one flat list of every question that speaks to Place Value, each one an equal vote. Note the Weight column — the midyear screener carries most of the answer.

The same body of evidence in a roll-up district, with questions grouped by sub-standard

The same evidence in a roll-up district. The questions are grouped by the sub-standard they belong to, and each group header carries the score that went into the average — the arithmetic the domain actually did.

Standards aligned above the leaves

Assessments don’t only align questions to the bottom of the tree. A question can be aligned to a cluster or a domain — the standards wheel above has one: a midyear task aligned to Solve, which is itself the parent of One Step Word Problems and Comparisons.

Read that alignment as a statement about coverage: this question speaks to Solve, and specifically to the part of Solve that the standards beneath it don’t already cover. A cluster stands for a set of ideas; its children carve named pieces out of that set; a question aligned to the cluster itself is evidence about the remainder — the complement of the children within the parent’s meaning.

It is not a shortcut for “this question covers everything under Solve.” If that is what you mean, align the question to each child instead — that alignment does reach them, because you said so explicitly.

Everything else follows from that reading, and it holds in both algorithms:

  • The children get nothing from it, because it isn’t about them. One Step Word Problems has no question of its own, so it stays Not Assessed — grey in the wheel — even though its parent has evidence and its sibling has a score. That grey is accurate rather than a hole in the data: nothing has assessed that idea yet.

  • It counts for every ancestor, like any other question. Solve’s question is part of Problem Solving and Perseverance, and of the grade as a whole.

  • In roll-up mode it is a sibling, not a summary. Since a cluster’s own questions cover a peer idea rather than a digest of the children, averaging them alongside the children is the arithmetic that matches the meaning. Solve has one child with evidence and one question of its own, so Solve is a 50/50 average of the two. If One Step Word Problems were assessed next month, Solve’s own question would settle to a third of the answer without anyone touching an alignment.

One thing to watch: the complement is only as real as the tree. If a cluster’s children were meant to name everything in it, a question aligned to the cluster has no leftover ideas of its own to speak to — but it will still be averaged as though it did. That’s usually a sign the question wants a more specific alignment, or the cluster wants another child.

Which one is my district using?

Roll-up is enabled per district by Forefront, and there is no district-level setting for it, so the quickest answer is to ask your Forefront contact. Two things in the app will also tell you:

  • Open a body of evidence and expand an assessment. If the questions are grouped under sub-standard headings with a score on each, you’re in a roll-up district. A single flat list of questions means the default algorithm.

  • Check the arithmetic on a wheel. In roll-up mode a standard is the average of the sub-standards drawn around it — skipping the grey ones, plus its own questions if it has any — so the rings should roughly add up. (Only roughly: the ring colors are rounded and the average is computed from unrounded values.) In the default mode the rings often won’t add up at all, because the value is an average of questions rather than of sub-standards.

Switching a district over rebuilds its stored proficiency, so numbers change on reports that nobody edited — worth knowing before a data meeting, and worth scheduling deliberately rather than mid-conference-week.

Did this answer your question?