MJD-1.0.0 explained
How the MathJewels Difficulty rating works
MJD describes what a problem asks a prepared learner to know, figure out, and carry through. It keeps those demands separate so a short, insight-heavy puzzle does not look the same as a long, routine exercise.
MJD is the name of the MathJewels Difficulty system. MJ is the prefix inside its compact labels.
The short version
The three public label fields
The full record tracks more evidence; the compact label keeps these three demands distinct.
What must already be known?
The content coordinate places the controlling skills in the MathJewels learning sequence.
How hard is the plan to find?
Challenge measures solution-finding demand while holding prerequisite mastery constant.
How much work remains?
Workload measures execution, writing, construction, and bookkeeping once the method is known.
01 · Read the label
Coordinates, not a total
The compact public label is built from structured rating fields. In MJ 4.3 · C2 · W1, each part has a different job:
- MJ 4.3 Content
- The required knowledge falls around point 4.3 in the K–8 learning progression—roughly early Grade 4 content. It is not a score out of ten.
- C2 Challenge
- A learner who has independently mastered the prerequisites faces standard problem-solving demand.
- W1 Workload
- After a method is known, only a light amount of calculation, drawing, or explanation remains.
Why separation matters
A problem can use early-grade content and still demand a C4 insight. Another can use later content but be C1 routine. A C4 problem may be elegantly short; a C2 problem may carry W3 bookkeeping. None of those combinations is contradictory.
Audience grade is separate, too. It describes whom a worksheet was designed for. The MJ coordinate describes the mathematical knowledge required by the easiest reasonable solution path.
02 · The comparison point
The reference student and controlling path
MJD needs a stable comparison point. Unless a problem names another population, it uses a reference student who has independent mastery of every listed prerequisite, typical reading fluency for the content band, no previous exposure to that exact problem, and only the materials and support the problem allows.
Hold readiness constant
Challenge is not raised because an unprepared student would struggle with prerequisite knowledge. It is not lowered because an expert sees an advanced shortcut.
Use the controlling path
The rating follows the easiest reasonable, prompt-compliant solution path available to that reference student—not an obscure trick and not the union of every possible method.
One rating belongs to one exact problem
The wording, diagrams, answer choices, response format, hints and scaffolds, calculator or manipulative rules, scratch-work expectations, materials, and time limit are all part of the rated item. A meaningful change creates a new revision and rating, because it may change the task.
For a combined multi-part item, content uses the highest required coordinate, challenge follows the hardest reasoning transition, and workload includes every required part. Independent parts should generally be rated separately; dependent parts stay together when later work materially relies on earlier results.
03 · Content
When should the required knowledge be available?
The content coordinate locates the latest skill milestone required by the controlling solution path. It answers a readiness question, not a question about how clever, long, or intimidating the problem feels.
K–8 sequence
K.0 8.9- 4.0 is approximately the beginning of Grade 4.
- 4.5 is approximately the middle of Grade 4.
- 4.9 is approximately the end of Grade 4.
The tenths are approximate positions in an instructional sequence. They are not calendar months, a child’s “grade level,” or claims of psychometric precision.
Beyond Grade 8, mathematics branches
The MJD design does not pretend that all advanced mathematics fits one ladder. It provides course coordinates such as HS:A1.4, HS:GEO.6, UG:CALC1.5, or GR:ALG.2. Genuinely research-level material can use a field coordinate such as R:COMB. Coordinates in different branches should not be compared as though one number made a universal ranking. These formats describe how the system can expand; the current worksheet collection occupies a much narrower content range.
Skills and mastery phases
Each target or prerequisite skill can have milestones for introduced, developing, independent, fluent, and transfer use. The grader identifies the mastery demanded by the actual problem, looks up the earliest matching milestone, and uses the latest controlling milestone across the chosen path.
Content can rise when…
- a later concept, theorem, representation, notation system, or proof technique is genuinely required;
- the prompt requires a particular advanced method;
- the controlling skill must be used at a later mastery phase.
Content does not rise merely because…
- numbers, prose, or calculations are long;
- the problem is a puzzle or hard to discover;
- the learner must be careful or manage many parts;
- the page looks advanced or intimidating.
Range and content confidence
A rating includes lower and upper content bounds when sequencing is uncertain, plus a confidence value for the placement. A visible estimate of MJ 3.3 may therefore have a plausible range around it. The range is more honest than pretending every skill has one exact universal teaching date.
04 · Challenge
How hard is it to find and organize a successful solution?
Challenge assumes the prerequisites are mastered. It measures planning, selection, reasoning, novelty, constraints, and justification—not the amount of arithmetic left after the plan is clear.
| Level | Core meaning | Typical signals |
|---|---|---|
| C0Tutorial | Recognition, imitation, or completion of a substantially modeled process. | The method is named or demonstrated; the learner fills in or copies a step. |
| C1Routine | One familiar, obvious procedure with almost no method choice. | Direct fact recall, an explicitly cued algorithm, or one familiar conversion. |
| C2Standard | Normal independent work for a learner who knows the content. | Choose a familiar method, make one or two decisions, or justify a standard result. |
| C3Stretch | A nontrivial plan built from familiar ideas. | Combine concepts, reason backward, organize cases, or manage interacting constraints. |
| C4Insight | A genuine key insight that is not substantially telegraphed is required. | Find hidden structure, an invariant, symmetry, construction, substitution, or useful reformulation. |
| C5Olympiad | Sustained original strategy with multiple significant insights or a difficult proof. | Build and connect a central chain of insights; “hard for the grade” alone is not enough. |
Standard is simply the name of the C2 challenge band. It does not mean standards-aligned.
Five evidence factors
Each factor is scored from 0 to 3 as evidence. The scores support a holistic decision; they are not added into a mechanical total.
- Method selection
- From an explicit method to an open strategic choice.
- Reasoning depth
- From an isolated step to a long dependency chain.
- Novelty
- From a rehearsed form to unfamiliar structure.
- Constraint management
- From no interacting constraints to several tightly coupled ones.
- Proof or justification
- From an answer only to sustained proof or validation.
Borderline cases default downward. C4 and C5 require a named key move or chain of insights; high workload by itself cannot produce an Insight or Olympiad rating.
05 · Workload
Once the method is known, how much execution remains?
Workload covers calculation, writing, construction, representation production, and bookkeeping. It deliberately separates “finding the idea” from “carrying the idea through.”
| Level | Core meaning | Time anchor |
|---|---|---|
| W0Negligible | One response or about one mechanical step. | Often under one minute after the method is known. |
| W1Light | A few calculations, a short diagram, or a brief explanation. | Commonly one to four minutes of total active time. |
| W2Moderate | Sustained ordinary calculation, several cases, or a meaningful explanation or construction. | Commonly four to twelve minutes of total active time. |
| W3Heavy | Lengthy or error-prone execution, extensive casework, or substantial written justification. | Commonly twelve to thirty minutes of total active time. |
| W4Extended | A long proof, investigation, project, or computation. | Normally more than thirty minutes of total active time. |
What supports the W level
- Mechanical execution
- How much routine calculation or symbolic work remains.
- Written output
- How much explanation, proof, labeling, or reporting is required.
- Representation production
- How much drawing, graphing, table-building, or construction is needed.
- Bookkeeping
- How much state, case, or intermediate-result tracking is required.
W level and active time are related, but not identical
The published estimate includes a minimum, typical, and maximum total active time for the reference student. The W bands use time as an anchor, but W is based primarily on execution after the method is known. A long search for one elegant insight can increase total time without creating heavy mechanical workload.
06 · Behind the label
How a rating is made
Canonicalize the exact item
Collect the statement, parts, assets, expected response, and all conditions. The system records an exact content hash.
Check gradeability
A clear item is gradeable. A small explicit assumption may produce gradeable-with-assumptions. Missing or contradictory material makes the item ungradeable rather than guessed at.
Solve and verify it
The grader establishes the expected answer when one exists and verifies at least one valid solution before judging difficulty.
Enumerate reasonable paths
Materially different methods are considered, including shortcuts and whether each path is realistically available to the reference student.
Select the controlling path and skills
The easiest reasonable prompt-compliant path controls. Target skills, prerequisites, mastery phases, and alternative-path skills are kept distinct.
Rate each dimension separately
Assign content with a range, choose C0–C5 holistically, estimate W0–W4 and active time, then assess presentation and item quality.
Add only verified external mappings
Curriculum crosswalks are supporting evidence. They never define the first-party MathJewels coordinate and are omitted when evidence is unavailable.
Record confidence, review, and version
The structured record is validated against the MJD schema. Published ratings are immutable; a correction creates a new rating that supersedes the old one.
Details kept outside the compact label
Presentation
Translation (T0–T3), language (L0–L3), and notation (N0–N3) help distinguish mathematical demand from how the task is communicated.
Assessment quality
Construct alignment, shortcut risk, guessing risk, ambiguity, and flags capture whether the item measures what it appears to intend.
Solution analysis
Expected answers, concise solution summaries, controlling-path decisions, common errors, and step estimates make the record auditable without publishing private reasoning.
External crosswalks
Verified mappings may locate the content in another curriculum or standard. They remain evidence, not an endorsement or replacement for MJD.
07 · Evidence strength
Confidence is not a child’s success probability
Confidence describes support for a placement or rating. It does not mean “the answer is probably correct,” and it does not predict whether a particular learner will succeed. Content confidence belongs specifically to the coordinate and its plausible range; calibration confidence describes the overall evidence supporting the rating and its crosswalks.
Calibration status tells you what kind of evidence exists
- Expert estimated
- Rubric judgment only.
- Externally crosswalked
- Rubric judgment supported by one or more verified external mappings.
- Pilot
- Limited student-response data exists.
- Calibrated
- Sufficient empirical response data supports the item estimate.
Where the current collection stands
Published MathJewels ratings currently begin as rubric-based estimates; some are strengthened by verified curriculum crosswalks. “Expert estimated” is a calibration-status name, not proof that a human expert has reviewed the item. Ratings should not be described as psychometrically calibrated unless sufficient connected student-response data supports that status.
When the rubric requires manual review
A rating must enter review when any trigger below applies. A trigger is a queueing rule; it does not by itself prove that review has been completed.
- The item is not fully gradeable.
- Content confidence is below 75%, or its K–8 range spans more than 0.6 grade-year.
- Challenge is C4 or C5.
- Automated graders differ by more than 0.3 grade-year, one C level, or one W level.
- A new or unmapped skill is needed, or crosswalk evidence materially conflicts.
- Shortcut, guessing, or ambiguity risk is 2–3, or construct alignment is weak.
- A generated item falls outside its approved family range.
- The item depends on undeclared specialized knowledge or involves high-stakes, safety, or legally regulated content.
Future empirical work should separate first attempts from hinted attempts, condition results on prerequisite mastery, preserve the exact item revision, and avoid treating raw percent-correct as intrinsic item difficulty.
08 · Practical use
Choose a worksheet with intention
Start with content readiness
Look for the required skills and a coordinate near content the learner can already use independently. Audience grade is a clue, not the rating itself.
Dial the kind of challenge
Choose C1–C2 for more direct practice, C3 for a nontrivial plan, or C4–C5 when the goal is insight and sustained strategy.
Budget the workload
Use W level and the active-time range to fit the available session. A low-W puzzle can still be deeply challenging.
Same content, very different challenge
Three worksheets in the current collection make the separation especially clear. All use the same MJ 2.7 content coordinate while asking for progressively deeper problem-solving.
Problem 9
Nurse Joy’s Poké Ball Packs
MJ 2.7 · C3 · W2Stretch challenge · Moderate workloadProblem 10
Rotom’s Pokédex Rectangle
MJ 2.7 · C4 · W2Insight challenge · Moderate workloadProblem 11
The Oran Berry Passing Circle
MJ 2.7 · C5 · W3Olympiad challenge · Heavy workloadA real label from the collection
Problem 10
Rotom’s Pokédex Rectangle
Content
MJ 2.7 is roughly late Grade 2 content. Its controlling skills are use hundred-chart structure and add and subtract quantities.
Challenge
C4 · Insight: the worksheet asks for a proof about a hundred-chart relationship whose useful structure is not substantially telegraphed. Discovering the path—not advanced content or long arithmetic—creates the C4 demand.
Workload
W2 · Moderate: Several coordinated calculations, representations, or explanation steps. Estimated active time is 7–18 minutes, typically 12 minutes.
The label describes the problem under its stated conditions. It does not label a learner, promise a teaching sequence, or say that everyone at the named content point will have the same experience.
09 · A separate system
Difficulty informs rewards, but it is not the reward
Jewel value is downstream of MJD and stored separately. The content coordinate selects the reference population for the time estimate; typical active time supplies the work estimate; and challenge adjusts for the collection’s assumed solve probability. Later content does not automatically earn more jewels.
Current reward policy
typical minutes ÷ (5 × challenge solve probability)Round up to keep expected work at 5 minutes or less per jewel, with a minimum of one.These percentages are an incentive-design policy, not empirical claims about a particular learner. Jewel allocations may change without changing the mathematical MJD rating.
10 · Common questions
Questions parents and teachers ask
Does MJ 4.3 mean “fourth grade, third month”?
No. It means an approximate point in a Grade 4 instructional sequence. The decimal is not a calendar month and is not a precise psychometric grade-equivalent score.
Can a Grade 2 content problem really be C4 or C5?
Yes. Content and challenge answer different questions. Elementary knowledge can support a very difficult insight or proof, just as advanced content can appear in a routine exercise.
Why can two similar-looking problems have different ratings?
A diagram, answer choices, required explanation, visible shortcut, hint, response format, or allowed material can change the easiest reasonable path, the challenge, or the workload. MJD rates the actual item, not its topic alone.
Why does a worksheet’s audience grade differ from its MJ coordinate?
Audience grade is an editorial choice about who may enjoy or benefit from the worksheet. The MJ coordinate comes from the knowledge required by the controlling solution path. Those can reasonably differ.
Is the active-time estimate a time limit?
No. It is an estimated range of focused work for the reference student. Real learners vary, and taking longer or shorter says nothing by itself about mathematical ability.
Does “expert estimated” mean a human expert approved the rating?
Not necessarily. It means the rating is a structured rubric estimate rather than an empirical calibration. The status name alone does not certify that manual review has been completed.
Will ratings change as MathJewels collects data?
The rubric and current rating remain versioned and auditable. New evidence may justify a new rating or calibration status, but published records are kept so historical attempts can still be interpreted.