Unrealistic Tests Might Be Our Last Hope for Education and Technical Interviews

The surprising value of the closed book, artificial tasks and trivia.
2026.10.04

Introduction

The following are understandable criticisms of schooling and testing (and not just because I have held them all at some point).

  1. Timed exams are unfair because the conditions are unrealistic.
  2. Closed-book exams are especially unrealistic.
  3. Curricula should emphasise more “real world content”.

These criticisms are commonly levied against high schools, universities, technical interviews for jobs, and many other domains. Their resolution lately seems to involve increasing use of assignments (to better reflect real world scenarios) and in turn embracing AI (because its use can barely be regulated otherwise).

So, the purpose of this post is to explore whether making education more realistic in the above sense is prudent. Specifically, whether these trends interfere with students’ learning of core knowledge and core skills (e.g., writing, low-level programming, mental arithmetic, memorisation, etc.).

In short, I will argue that the relationship of learning the above to learning from projects is like the relationship of exercise to playing a sport. Moreover, trying to learn with an unexercised mind is especially difficult, and exercise is often best done in unrealistic settings. Furthermore, testing ought to assess how exercised a mind is, and this is best done AI-free or with certain closed-book components such as oral examination or on-the-spot tasks.

In the context of technical interviews, my argument is similarly that we should not shy away from trivia, i.e., in place of more realistic programming assignments. By “trivia” I mean being peripheral, perhaps strange or edge-case questions, whose answers tend to be learned through experience, interest, getting one’s hands dirty, deep rather than surface level reading, etc.

Of course there are a lot of wrong ways to do old-fashioned education and testing, but I hope to demonstrate there is actually a lot to be gained by taking a step back and figuring out how to educate and test effectively, rather than clumsily indulging AI inevitabilism.

Why "Realistic" is Impossible

To give the above criticisms their due, I will begin with an affirming example from the Mathematics coursework I did in the mid 2010’s. In my bachelor’s degree, I hated the exams; I hated the time pressure, I hated the memorisation. I thought “when am I ever going to have to solve these problems under time pressure?”, “when am I not going to be able to look things up to help?, ”I should be getting tested on more realistic problem-solving skills”, and so on.

Then, in my master’s degree, I was introduced to an alternative I loved: take-home exams. These were harder assignments, to be completed over a 2-3 day window, i.e., more quickly than usual. Most, importantly, even being open-book (and open-internet), I found that these exams could not generally be finished in 2-3 days unless one had thoroughly studied the content beforehand.

However, with generative AI, those kind of assignments / take-home exams are simply infeasible because they can now be completed with minimal understanding, i.e., compared to a timed, supervised, closed-book exam (or take-home exam of years past).

This point was strikingly demonstrated in early 2026 when Professor Roberto Serrano of Brown University had to turn a midterm exam for ECON1170 (“Welfare Economics and Social Choice Theory”) into a take-home exam. The students almost all attained perfect grades for it, and yet in the subsequent in-person final exam, they mostly performed much worse (see Figure 1 below from the original article). This suggests that their success on the take-home midterms was achieved with minimal understanding of the underlying content.

Figure 1: ECON1170 take-home midterm exam grades (orange) vs in-person final exam grades (grey).

Similarly, HTMX’s creator and Montana State University instructor Carson Gross said in an interview that because his students now typically score all A’s in assignments due to AI, he is increasing the weighting of his pen-and-paper quizzes to 50% of their total grade. Previously, ~70% of their total grade came from a large programming assignment.

However, in his new scheme he is not merely cynically introducing a mechanism for balancing grades, these quizzes have also been modified to emphasise annotation, review and analysis and thus represent deeply valuable fundamentals (more on this later).

So, before getting into other strategies for testing without overreliance on assignments and take-home exams, I want to discuss what it means to educate with an emphasis on fundamentals, and why it often needs to be done in unrealistic settings.

Leaning Into the Unrealistic in Sports and Language Learning

The relationship between unrealistic to realistic learning, is what exercise is to sports. For example, it is not controversial to task a basketball player with bench presses and tricep pulldowns to put them in a better position to develop their jump-shooting skills. The fact that they would never perform those specific exercise-movements in a real basketball game (let alone see a workout bench) is a feature, not a bug.

Moreover, that basketball player will perform all sorts of “drills” to develop specific skills like dribbling, or decision-making involving one, two or three other players rather than the full nine other players in a real game. We do not think of exercise or drills as “unrealistic”, we think of them as a developing a player’s foundation; a controlled environment to precisely reflect on and develop skills.

In fact, the basketball world is already grappling with the problem of “too much realism”. Specifically, in the US, top high school prospects play more games than ever (especially via constant ad-hoc AAU tournaments), and veteran players (e.g., the late Kobe Bryant) argue that this has come at the expense of those players spending time being meaningfully coached and practising their fundamentals.

Another example of this dynamic is revealed by a common complaint by language learners (especially those using textbooks and apps like Duolingo). Specifically, learners complain about having to translate unrealistic sentences, e.g., “the elephant apologised to the shopkeeper for stealing three green race-cars”.

While this is a valid complaint if you just want to prepare for a two-week holiday in another country, I believe the complaint is actually shortsighted when it comes to long term language learning. The reason for this is that translating realistic sentences can often be done in lieu of fundamental language understanding, i.e., by simply using context or process of elimination. Whereas translating a truly strange sentence forces one to really understand from first principles why that strange sentence is so.

The Challenge of Learning with an Unexercised Mind

The same way learning a sport is harder with an unexercised body, so is much else with an unexercised mind.

I once saw this first hand 8 years ago when I tutored a family friend who was struggling with learning algebra (specifically, factorising quadratic equations). Her issue was that she struggled to do basic arithmetic without a calculator, and despite her having a calculator, this was a serious obstacle to her learning algebra.

At the time, a popular opinion would have gone: “mental arithmetic is no longer an important skill; let students use a calculator and focus their efforts on higher order, conceptual topics like algebra”. However, the constant context-switching between her trying to learn algebra and her constantly reaching for a calculator seemed to be really stunt her learning.1

I have tried to apply this lesson to my own life by continuing to practice coding by hand, and in particular solving problems without AI. Specifically, thinking problems through and only looking up facts in documentation. The point is not just to maintain an exercised mind, but to also develop the skill of judging when outsourcing one’s thinking is actually worth it – a skill that requires knowing how quickly and effectively an exercised mind can solve a problem.

A related question to consider is: can one learn the interesting stuff by outsourcing all the boring stuff? For instance, can one be a great storyteller without being great at articulating (e.g., great at writing, speaking, performing, etc). Likewise, can one conceive of good arguments without the ability to make good arguments.

The Value of Unrealistic Curicula

Without strong fundamentals, learners often lack the knowledge of key questions to ask of their own work. For instance, many computer science programs have deemphasised fundamental topics like concurrency, operating systems and networking in favour of exposing students to more “real world” tech stacks. Arguably, this leads engineers to struggle to even conceive of what performance bottlenecks are in their purview (Casey Muratori puts it well here). The above deemphasised topics are considered “unrealistic” because engineers typically rarely work directly at those levels.

I started to think about this more when delivering lectures on Turing machines in a “models of computation” subject in 2025. Specifically, I wanted to convey to students why the subject was relevant. Moreover, I wanted to do this without concocting nebulous real world applications, while also appreciating that only some students would care about the unit’s theoretical insights e.g., the Church-Turing Thesis or undecidability results. Thus, I made the case that a good software engineer should have the kind of mind that can program at the level of a Turing Machine, and moreover a computer science degree should signal that you have that kind of mind.

On a related note, the drive to front-load industry trends via universities makes me think we should actually be suspicious of educational institutions being held responsible for all knowledge and opinions, i.e., perhaps some things should be decided on the ground. For instance, I was listening to an Interview of divorce attorney James Sexton, where James laments that schooling had taught him a lot about dividing fractions, but nothing about conflict resolution. I believe he meant well, but is it really a school’s job to set parameters and opinions on communication between lovers? Is that not the job of families and communities? Indeed, those sorts of sentiments are among the many calls for “the realistic” in education that I think have gone too far.

How Else Educators Can Test

Thus, if we accept that an exercised mind is a must and that unrealistic settings and content may be required – what can educators do?

Here are some ideas:

Technical Interviews and the Value of Trivia

I first started to think about this concept when I came across the YouTuber Coding Jesus, who records mock technical interviews for programmers. A common criticism he receives is that he makes the interviewees look unqualified by asking them what are basically “trivia” questions about programming and software.

To the contrary, I find that the brightest candidates in his videos are reliably the ones that get most of the trivia right and otherwise seem interested in the questions that they get wrong (see Footnote 2 for an in depth analysis of one interview). The idea is that knowing well-chosen programming trivia can indicate that a candidate has good exposure to reading other people’s code; that they have read articles deeply rather than by summaries only; and most importantly, that they have gotten their hands dirty and stumbled upon trivia, likely through troubleshooting problems.

Ultimately, I think our aversion to trivia comes from an obsession with testing being “fair” in the sense that it is “predictable”, “free of chance” and free of “serendipity”. In other words, the task should be so clear as to guarantee perfect scores given enough mechanical effort.

Perhaps the point of testing is not to find the examinees who get every question right.

On one hand, too much predictability can lead to grade inflation, which arguably makes everyone unhappy. On the other hand, it is often self-defeating. For example, this year I gave an “exam preparation and practice questions” lecture for a web-development course. At the end of the lecture, I was swarmed by students demanding to know more and more about what would be in the exam. Ironically, I told them, the more we tell you about what will be in the exam, the more “tricks” we will have to insert into the exam just to differentiate you.

Conclusion

The best education is about developing a student’s mind, and testing should verify the extent to which that has been achieved at a particular point in time.3 Moreover, testing can be divided into verification while learning is happening (i.e., during education) and after the fact (e.g., in technical interviews).

While it seems desirable to educate and test via large scale projects, completed in realistic working conditions, it is unfortunately misguided to do so nowadays. On one hand, this is because AI assistance can enable completion of those projects at the expense of a student’s understanding. On the other hand, students need controlled, distraction-free environments to hone in one and to develop specific skills.

Instead of taking our feet off the brakes and risk letting the value of education deteriorate, I believe we need to tune our existing brakes for the new surfaces we’re driving on.

Ultimately, pen-and-paper or offline computing should be the predominant medium, and subjects emphasising fundamentals skills should predominate courses. Then, insofar as projects should be used, human-centric and improvisational add-ons should be an increasingly large part of those projects (e.g., oral examination).

Then, in the post-educational context (e.g., technical interviews), testing should capture signals about whether a candidate is deeply interested and curious, have actually gotten their hands dirty, and are not shy about reading deeply. Importantly, to achieve that, subtle and strange questions often need to be asked (i.e., trivia), and most importantly, we ought to be ok with that process making predictability and perfection hard to come by.

Tags: EducationTechnical InterviewsArtificial Intelligence

Comments

Comments are a static snapshot of a GitHub Issue. Please leave a comment and after reviewing it, I'll rebuild the site with it.

Footnotes and References


1

Another example of this I constantly hear these days is specifically about why we must learn long-division.

Again, I am confused by those same people not surmising that learning long-division is not about learning to divide by multi-digit numbers, but to learn how to understand and follow algorithms. Moreover, the reason that using the unrealistic example of long-division makes sense, is because it does not require much further background knowledge.

2

The interviews that comes to mind are two that he did for python (a less common topic on his channel, but one which I know best).

In the first (beginner) one (that went poorly), the questions start with:

  1. How do you check your python version (in particular, how to do it programmatically)?
  2. What is a REPL?
  3. What is type hinting?
  4. What is integer division?
  5. Where in memory (i.e., heap vs stack) is a class stored?
  6. How do you create an empty class?
  7. What is a difference between a tuple and a list?

To reiterate, the point is not to get them all right, but a good candidate should know most of the answers.

I think it is fair to argue that the above questions are trivia as far as beginner level knowledge goes. However, I think being able to answer those questions reveals the following signals.

  • Questions 3 and 6 signal that the candidate has looked into other people’s code (esp. documentation and open source), where type hinting and empty (abstract base) classes are common, i.e., as opposed to LLM code outputs. Similarly, guides using REPLs are more common in hand-written documentation than LLM outputs.
  • Question 1 signals that the candidate does not just run python on perfectly pre-calibrated google colab servers. Tougher variants about dependency versions and packaging would imply that the candidate has gotten their hands dirty due to the inevitable dependency hell python developers reach.
  • Similarly, an answer to Question 7 (usually: “tuples are immutable”) is usually learned by playing around with tuples and causing an exception to be raised.
  • Answers to Question 5 (e.g., heap memory for the class, but stack memory for the reference) are rarely explicitly taught in python courses, but being able to figure an answer out implies good fundamentals in operating systems.

In the second (advanced) one (that went well), you will find a tougher variety of questions concerning object ids, iterators, the GIL, etc. You can just tell the candidate is bright by the way he answers those questions, I’m not so sure his brightness would have come through by just being asked to invert a binary tree live.

3

Although, one weakness pervasive in testing (whether the questions are well-designed or not) is that grading typically does not reward improvement as well as it rewards students being comfortable from the very start. Perhaps I’ll look into that more, one day.