Building the global standard for measuring and Credentialing AI capability @answerr_ai MBA'25 @babson

Boston
Don’t give up. Be dead instead.
1
12
1,159
We put our AI in front of 211 students at the biggest entrepreneurship school in the world for a full term. One number caught me completely off guard. Going in, I expected to see usage go up. What surprised me was how much more curious the students became. They started asking bigger and stranger questions. Curiosity-driven prompting rose more than 40% over the term. And there's a simple reason for it. When you remove a hurdle from someone's day, they explore more. Funny how that works. Run a university program and you can see this same picture for your own students.
8
14
374
Mohd Qaiser Malik retweeted
Everyone around me wanted the safe, standard path. Throughout my entire life, I have deliberately chosen the opposite. • I worked on multiple projects with the Government of India. • Then I opened an IT services company I ran for 4 years • That paid for my move to the US, where I did my MBA at the largest entrepreneurship school in the world. And that turned into, @AnswerrAI. We're going to have an impact on millions of lives. So much more to come.
2
15
1,264
Mohd Qaiser Malik retweeted
There's no such thing as an “AI skill”. And it's the biggest reason hiring for AI capabilities feels a little off right now. Every job post wants "AI skills" or "prompt engineering." However, prompting isn't necessarily a skill you can hold onto. The models get better and better at reading intent every month, so whatever you test for today will inevitably be gone within the next few months. Almost anyone can type a question into Claude and paste the answer back. An since everyone can prompt, the validity of prompting as a way to measure skill, lessens. What’s actually scarce is 1) complex reasoning, and 2) the judgment to catch a model when it's clearly wrong (but sounds right). In 2026, you can’t be accepting the first “clean-looking” answer. If I were writing a job req today, I actually would not ask whether someone can use AI. I'd sit them in front of it for an hour and watch what they do when it hands them a confident, wrong answer. Do they catch it, or forward it on and move to the next thing? That’ll tell you more than anything.
4
2
15
1,183
Mohd Qaiser Malik retweeted
2027 prediction: The companies who are "fastest" to move with AI, will completely collapse. When your team moves "fast" with AI, that usually means they're rushing. They'll grab the first answer AI gives them and run with it, without ever stopping to check whether it was correct or not. And you can watch this happen in real time, too. The person being driven by the AI: • takes the first answer as the final answer • doesn't catch the model when it's confidently wrong • can tell you the answer appeared, but not why it holds up • never asks a second, sharper question Someone actually using AI correctly assumes the first answer can be better, so they push back and keep working until it holds. They treat the model like more of a "sharp intern" they still have to manage. The problem is those four habits never show up on a usage dashboard. You see them after a wrong answer is already out the door. We measure them directly in a 60-minute assessment. Want to see what that speed is costing you? Schedule a time with our team here: answerr.ai/
2
2
16
1,263
Mohd Qaiser Malik retweeted
We built a startup to measure anyone’s capability with AI in 60 minutes. Most tools measure how much someone uses AI or whether the output looks polished. However, that tells you almost nothing about the person. What tells you something is: • how they handle the model when it's wrong but convincing • what happens when it hands them something plausible they can't check? • do they notice, and how fast? So with @AnswerrAI, we measure the whole conversation. There are five things we benchmark on: 1. Judgment Quality: knowing both A) when to trust the model and B) when to stop handing it more work. 2. Question Originality: asking something better than what the model has answered a thousand times already. 3. Critical Thinking: pushing on an answer that sounds right but hasn't been verified. 4. Complex Reasoning: holding a hard problem together and dismissing the urge to take the first “clean-looking” output. 5. Learning Velocity: whether someone gets sharper over a term or not. The whole thing is behavioral, so there's nothing for students to cram for. It watches what you do when you don't already know the answer. This way, a dean gets an actual benchmark to measure a student against a semester later.And a hiring manager gets to see how a candidate thinks before they make an offer to them. Go here and I'll run one for your team or class: answerr.ai/enterprise/capabi…
2
14
819
Mohd Qaiser Malik retweeted
Every university professor I demo @AnswerrAI to asks me some version of: "so it spies on my students?" It’s actually the other way around. Right now, teachers are the ones who can't see what their students are doing. Think about where learning actually happens now with assignments: • A student opens ChatGPT • They blindly work through the whole assignment in there. • The next morning they paste a clean answer into the portal. • The teacher sees that answer, and they never see the thinking that did / didn't happen in order to complete the assignment. The whole process is invisible to them. And that is the ultimate goal of @AnswerrAI. To help teachers finally see 1) who's thinking and 2) who's just pasting answers in. This way, you’ll know which students will actually be prepared to enter the modern workforce, where the ability to effectively use AI will be invaluable. (and no, before anyone asks, we don't train any model on their students' work) Interested in a demo? Send me a message and we’ll find a time.
2
10
1,200
Mohd Qaiser Malik retweeted
My co-founder and I are building the standard to measure AI capability. Some days this work is exciting. Other days it's draining. I'm building something I believe will reach millions of people, so you'd think the hard part is the technology. It isn't (not anymore at least). Building has never been easier. Models write most of the code, and anyone with an idea and a weekend can ship something. Getting people to actually use what you built is the hard part today. You can build the best product in your category and still have a failing product because nobody knew it existed. So my approach right now building @AnswerrAI: • Talk to 10 customers every week. • Try to understand what’s actually wrong • Learn what we can do to make the product better Distribution is the founder's job. You find your own distribution, or there is nothing to build. I’m extremely excited with the progress we’re making, and can’t wait to see where we’re at in the next few months!
4
2
17
1,180
Mohd Qaiser Malik retweeted
Dozens of universities have asked me: “what AI detector should we purchase?” In my opinion, none of them are worth it. Those tools are solving the wrong problem. Your students will use AI whether you catch them or not, and the jobs they graduate into will demand it. So the useful thing is to make them good at it. Think about hiring ten years ago. If you wanted a job at Salesforce, you had to actually know Salesforce, because the whole company ran on it. Nobody sold universities a Salesforce detector. They taught the tool and hired the folks who were genuinely fluent in it. AI is that now, for almost every job, from a marketing desk to a hospital floor. Detection only answers one (irrelevant) question: “did a machine touch this?” What matters is whether the student can push back on the answer and make a good call with it.
1
2
12
1,315
Mohd Qaiser Malik retweeted
How well someone works with AI is now a measurable number. However, most tools are measuring the wrong thing: How much they use it. This tells you nothing about the person. Usage counts how many prompts someone sent. It can't see whether they caught the model when it was confidently wrong, or just pasted the first answer and moved on. So, instead of reading the finished work, we started measuring the process. AIQ measures five behavioral dimensions that are relevant in 2 places: Hiring: Every resume is AI-written now and the take-home was most likely also completed by a model, so your interviews can't see how a candidate really works with AI. A 60-minute assessment gives you an accurate read before you 1) make an offer, and 2) spend $100k on one bad hire. Education: A transcript shows what a student turned in, but it doesn’t reveal whether or not they got better / more capable over a term. Our suite plugs into the LLMs they already use and gives the university a capability number that is measurable, so they can watch a student grow OR catch one who's falling behind. TL;DR: Adoption was last year's question. Now, capability is the most important factor. Book a demo for an AIQ report: answerr.ai/
2
15
1,106
Mohd Qaiser Malik retweeted
The most underrated thing you can do with AI is ask a sharper question. Everyone can pull an answer now. What separates people is the question they bring to it. The question you ask sets the ceiling on the answer you can get back. It's one of the five things we measure with @AnswerrAI, and we call it “Question Originality”. We look at the questions a person asks the model, before we ever judge what it gave back. Ask better questions than the next person, and you'll out-think them with the same tool open in front of you both.
2
16
1,368
Mohd Qaiser Malik retweeted
Nobody spends a hundred thousand dollars at a top university to learn. They spend it to get employed. And as it stands, employers are checking for something that the modern college degree isn’t measuring. I've spent most of this year talking to education institutes, around ten a week. Things move slowly there. A university will talk about a change for eighteen months before anyone implements it. This one won't wait eighteen months, because the pressure is coming from the hiring side. An employer filters the resumes with an ATS and spends the interview budget. They still can't see the skill they're actually hiring for: whether this person can catch a confidently wrong AI model and make a rational decision with it, rather than pasting every word it outputs. So with @AnswerrAI, we measure students the way an employer eventually will. How effective are they when working with LLMS? Do they verify what it hands them? Are they still able to make a good decision when it's wrong? The same assessment happens on both sides, actually. A university can run it across a semester, so every student has a baseline and a number that will actually improve / get worse over time. An employer can run it once, as a sixty-minute assessment, before a few bad hires costs them $1M+. Now that everyone has access to the same tools, you can expect employers to start measuring how effectively people can actually use them.
2
15
1,084