Check out the conversation on Apple, Spotify, and YouTube.
(0:00)
Aakash: What’s going to happen to my job?
Mikhail: What actually will happen is that teams will be smaller, leaner, and faster. Not only is it going to change, the lines are going to blur.
Aakash: Meet Mikhail Shcheglov, the CPO at OLX Classifieds, formerly CPO at Azerbaijan’s biggest e-commerce marketplace and GPM at Bolt.
Aakash: This is the coolest visualization I’ve ever seen. So what tool is this?
Mikhail: It understands our industry, classifieds, our market, our business model specifically at the level of 54%.
Aakash: What’s your process these days to hire PMs? How do you find PMs that are at this level of AI native?
Mikhail: I typically look at three things. Fundamentals, problem solving, systematic thinking.
Aakash: On top of our own job, we need to keep the pipeline of really good PMs. How can this help with recruiting?
Mikhail: This automates probably like 70 to 75% of the recruiting workflow.
Aakash: What’s going to happen to the future? Is product management going to exist in the future?
Mikhail: I have good news for you.
Housekeeping and bundle (0:54)
[Aakash asks listeners to subscribe on YouTube and follow on Apple and Spotify, then plugs the Product Growth Bundle at https://bundle.aakashg.com/, which includes a year of paid plans for Bolt.new, Airtable, Speechify, Descript, Magic Patterns, Linear, Dovetail, Arize, and Mobbin.]
Why this episode matters (1:30)
Aakash: I know you guys are overwhelmed with guides on Hermes and OpenClaw and Claude Code and ChatGPT. It seems like there are a million AI tools out there. And more and more of you are reporting to me that you’re not even feeling more productive after using AI.
So I have been searching for those PMs and product leaders who are not just getting a 10% boost in productivity but are getting a 2 to 3x boost in productivity. And Mikhail is one of those people.
What he has built at OLX Classifieds is an entire operating system. If you’ve been hearing about buzzwords like OpenClaw and Hermes and saying there’s no value for PMs, this episode is the one that will change your mind. He has got his entire team hooked into this operating system that he has built and they are validating features, shipping features, reviewing features. The possibilities for a PM are endless and I can’t wait for him to share this all with you.
Mikhail, welcome to the podcast.
Mikhail: Glad to be here.
The knowledge bottleneck problem (2:29)
Aakash: So how do people really build a company operating system? How do they measure it? What are the keys to getting the most out of AI as a company?
Mikhail: If you think about it, the core goal of having AI native teams is to be able to automate as much knowledge as we can. In the previous paradigm, what I call the old paradigm, you used to have a knowledge worker, an employee who has built context around a certain domain. Let’s say it’s a deep expertise domain like search, recommendations, ML, and stuff like that. And then all of a sudden this person decides to leave the company and takes all the knowledge with him or with her.
He is a bottleneck effectively, a knowledge bottleneck. And the core theme that we’re trying to prevent here is the leakage of that knowledge. Because if you could have a single storage of this entire business, customer, product, and technical knowledge in one place, then it would increase the entire value of your organization. That’s just one angle to it.
And another angle is the better your AI knows your context, the more autonomy you can give it, and the higher level tasks you can delegate to it.
Inside the company knowledge graph (3:55)
Mikhail: Let me explain how it actually works in our case. So what you see here is the entire knowledge graph of our company that has been formed over five months. Five months is basically nothing and you can already see this web of interconnected particles. And these particles cover everything, from our contacts, who talks to whom, to which projects we’re working on, to which customers we’ve talked with, which businesses we’ve reached out to, even our funnel metrics. Basically everything.
And you can see in the top right corner there is a specific metric which is product context coverage. This is one of my personal KPIs, because the higher this percentage of coverage, the more, as I mentioned, you can delegate to it.
Right now what this means is that it understands our industry, classifieds, our market, our business model specifically, the top line, the bottom line, the drivers behind it, our customers and the value drivers for them, at the level of 54%. Which already means that it can operate as a capable, I would say, junior to mid product manager, and even make backlog decisions.
The more this coverage grows, at numbers of 70 to maybe 90%, I would not be surprised if this knowledge graph would enable you to do some strategy level work and help the business navigate decisions. It’s a question of when we’re going to get there, but we’ll definitely get there.
Let me zoom in a little bit on the elements of this knowledge graph. You have three layers. The first layer is the product, so all the nodes that are connected to the product side of things, and you can see product has been very involved in everything.
You have the contacts, which is basically all the people talking to each other in different parts of the organization. And there are personal things that I share, my thoughts, my reflections, and the personal things that some of the other people share, fully anonymized of course.
Then another layer is the teams. I don’t really go at the level of actual cross-functional teams. It’s more of clusters or divisions. So if you take buyer and seller experience, you can immediately see how interconnected they are within the organization, or the platform teams, or the pay & ship teams. And you can even drill down at the level of a specific product manager and you can see on a team level what their knowledge of their existing context is, which for me is a very strong predictor of how good the team is in terms of customer discovery.
Not only that, there is another angle to it. If the team isn’t really well connected, and you can clearly see that this team has very little overlapping nodes with buyer and seller experience even though platform is a horizontal layer and ideally they should be intertwined with each other, you can clearly see that there are some instances of silos happening. Which means, one, teams should talk more with each other. And two, I would also raise the question of stakeholder management. Are stakeholders really being involved in the processes within the teams and so on.
That’s my core tool that I’m using, because it shows me the quality of the output of the entire organization.
How to measure product context coverage (8:00)
Aakash: This is the coolest visualization I’ve ever seen. So what tool is this? How are these metrics working? How are they measuring things?
Mikhail: You can use tools like Obsidian for this. I personally created this one from scratch using Fable, because it was a two-liner prompt basically.
The important question that you asked is a super important one, which is how do we measure this. We have a very specific prompt that asks an AI, what percentage of knowledge around industry, and you have to be very specific. What do you imply by industry? In our case we have five verticals. We have real estate, we have auto, we have goods sales, we have services, and jobs.
What percentage of knowledge around the business, meaning the business model, the actual P&L, the drivers. And what percentage around customers, like different buyer, seller, customer segments, even on a cohort level, marketing and so on. Which percentage of knowledge do you have at the current point in time, taking all of your memories loaded into you directly.
And what we’ve seen is that AI has been quite accurate in digitizing this abstract request into a specific number as an output. What I’m also observing is that the more active product managers are, and they typically come with transcripts or research studies or RFD documents and they load it directly into our agent, into its memory, I’m seeing that the metric is actually improving over time.
It doesn’t mean that it’s a single source of truth, and I think each organization should come up with a specific context metric that would be highly specific and relevant to them. But this is something that I found directionally works.
Sponsor break (9:57)
[Sponsor reads for Bolt.new and Customer.io. Full details and links are in the YouTube description.]
What the agentic CPO actually owns (12:08)
Aakash: Okay, so this is a totally different way to view a product team. Most product teams, most PMs aren’t used to this. If they want to build this, it sounds like there’s two things they do. They define the metric and then they create the visualization with Fable.
Let’s take a step back here. What is the role of a CPO in an organization like this? What is the new agentic CPO doing here?
Mikhail: That’s a great question, Aakash. What has changed significantly is that in the past the whole team was responsible for outcomes and outputs. Outcomes were measured in how much your metric has moved, and outputs in how many PRD documents, how much code, how many prototypes, Figma mockups your team has been able to produce depending on the function.
Right now what has changed, and changed pretty significantly, is that the cost of this output is negligible. But what becomes more important is the quality and the token consumption.
So what do I mean by quality? Quality for me is the amount of AI outputs that have significantly moved the metric. So it’s essentially outcomes, but outcomes that were driven by AI outputs. And this is how I measure it in terms of the metrics.
In terms of the actual tooling, I think previously what a CPO was responsible for is building procedural scaffolding on top of the organization. What it means is that you have certain hiring principles, your vision. Then if you go a level lower you have certain processes like a planning cadence, a yearly strategy, then a quarterly strategy, then sprint planning. And then you have different reviews like design reviews, product reviews and so on. And this allowed the organization to function.
But right now what is more important is not the process itself, it’s creating an operating system within which AI could be a collaborator of product managers. And where AI could not only help to come up with or challenge ideas, but could autonomously make decisions.
The best way to achieve this is by owning the entire agentic scaffolding architecture, creating specific rituals, creating specific processes, and enabling the team to be in touch with those agents all the time. Training the team to use it, actually.
Where the PM time savings come from (15:10)
Aakash: How does this manifest in the PM process? Where are the savings had for teams? Why should a CPO suddenly be spending so much time as an agentic orchestrator?
Mikhail: The biggest reason for this is, if you break down the time that product managers on average used to spend, and I estimated this based on my previous experiences, around 50% of the time of a typical product manager was spent on processes and rituals.
You have different business reports, weekly reports, stakeholder management reports, different demos and so on. And this is a repetitive type of task that doesn’t really require much cognitive input in many cases. It’s just something that you have to do manually.
If we abstract all of this and start delegating this part of work to AI, then essentially what you end up with is a product manager that is solely focused on discovery, where the biggest leverage actually is. And a single product manager could actually do the work of two product managers, because he doesn’t have the burden of all the processes that the typical corporate environment imposes on him.
The agent in Slack and status reporting (16:44)
Aakash: Wow. Okay. So we’ve seen the knowledge graph. We understand the role of the CPO and what it does for a PM. Can you show us what this agent looks like in practice? How a stakeholder might interact with it?
Mikhail: There are multiple buckets of tools that this agent can solve, and we constantly improve it. So the first bucket is status reporting. You can ask it, what’s the status of a specific project. Let’s say, what’s the status of auto spare parts catalog integration. Respond in English, because it might get finicky.
And while we’re waiting for this, there are multiple other things. Status can be, you can get a response in whichever shape or form you want, directly in Slack, in Google Docs, in Confluence, whichever format is more convenient for you.
Second thing is it is integrated into all of the workspaces. Google Workspace, meaning calendar, Gmail and so on. I don’t read my email anymore. An agent does it for me and pings me in case it’s urgent. I don’t manage my calendar anymore, and it’s especially difficult when we live in this interconnected world where everybody is remote in different parts of the globe and just setting a single meeting is hard.
Oh, you can clearly see it’s all there. Scope, listings, automated. You can clearly see what’s the status. It’s quite straightforward.
Aakash: Nice. And have you been tweaking the prompts on how it gives status updates, because this is a pretty good one.
Mikhail: Yeah. It has a massive library of rules and imperatives. I think like 700 lines or so, to give you the cleanest and the most factual output without hallucinations.
Training stakeholders to talk to the agent first (18:44)
Aakash: Very cool. So we’ll go into how people build that. What other use cases should they know about within Slack that they can accomplish here?
Mikhail: One angle to this. Let’s say you’re a stakeholder and you have a certain feature that you just assume is important to have for whatever reason. A typical route for a stakeholder in this case is to go directly to a product manager and harass a product manager with feature requests.
But now we have a gatekeeper, which is our agent running under the hood, and we train our stakeholders to interact with an agent before reaching out to a product manager.
Let’s say, yeah, I just woke up and I have a feature request. I urgently want to have a video submission for our stories, like an Instagram, because it’s fancy, because I saw somebody else doing it.
Aakash: That’s always the reason. All competitors are doing it. We are behind.
Mikhail: Yeah. For sellers. And this is just a typical feature request which doesn’t have a problem framing. It doesn’t have any impact estimates. It doesn’t have any rationale behind it and it doesn’t respect the actual priorities that have already been defined for the organization.
We’ll just give it a little bit of time to think. It will respond with a list of clarifying questions. And in case after answering all the clarifying questions an agent decides that this feature isn’t worth building, then it will politely say so. If it is worth building, then this feature will be added to the backlog and escalated to the responsible product manager. An agent has all the organizational structure mapped out internally, so it knows exactly who the owner of a specific domain is.
Why the CPO should own the agent (21:06)
Aakash: Wow. And do you consider yourself the owner of this Saul Goodman agent? A lot of people want to outsource that to an AI ops role or something like that.
Mikhail: I think it’s a great point and it’s a dangerous way to go. You should own this for two reasons.
One reason is the speed of iteration. If you own this and you’re constantly improving it, I’m getting feedback every single day on what works, what doesn’t work, and so on. And I immediately open my IDE or my Claude Code and I make the changes right off the bat as we go. It’s super fast from feedback to deployment.
Secondly, this actually has an impact on the organization, because one, it saves time, and two, it steers decision making, which means it might have an impact on the entire business. And I think the CPO should own this. If you delegate this to an engineering team or if you delegate it to another person who doesn’t have skin in the game, then you’re not going to get the same level of speed and impact.
The architecture, OpenClaw plus Hermes (22:14)
Aakash: Okay. So CPOs should be building this now. They want to know how. Can we open up the covers and show people how they build something like this?
Mikhail: Absolutely. Let me open up my IDE. I will show you two ways. This is just to give an overview. I will walk you through the architecture, the brain, the memory, the tools, the skills, and I’ll explain it all as we go.
For the architecture here what we’re using is a combination of OpenClaw and Hermes. Why those two? Because OpenClaw is just a great scaffolding. It has everything available out of the box and great engineering support.
But Hermes has a unique feature to it, which is automated skill generation. I’ve actually tested it across five core topics that my team works on and I’ve seen an improvement in a recall metric, which means that it was way more accurate in responding using those skills, by 31%. Which was a dramatic improvement.
So I decided to blend both of those scaffoldings together in a combination. And it’s jacked in terms of the capabilities, because we’ve been improving it for like five months continuously.
Three layers of memory (23:38)
Mikhail: When it comes to memory, it has three layers of memory.
So the first layer is what you’ve seen, a knowledge graph. It’s like an entire universe of interconnected knowledge and everything, and you can see the code here and so on.
But what is plugged in there? It’s a second layer, which is a vector database. The vector database is essentially every single piece of knowledge that an agent receives gets immediately converted into a vector. And you need this because most of the requests are fuzzy. In order for you to get good retrieval, you do need to have vectors, because you might ask some random stuff that is not related to any specific keyword. And an agent needs to be able to match your request to a specific point, to a specific number in your knowledge graph.
And the third layer, and what I’ve seen makes agents the most robust so they don’t lose context, is every single conversation is stored in transcripts. You can clearly see that every single day the agent stores everything into MD files. And not only does this encompass Granola transcripts, because all of our meetings are transcribed, but also all of the conversations that I had with an agent and all of the reflections that an agent has.
So those three layers of memories, persistent memories, they create this robust knowledge base which improves every single day.
Why summarization hurts recall (25:18)
Aakash: Wow. How do you architect that properly? So obviously you get a Granola enterprise license. You have that coming in and recording every single meeting. But I guess you don’t want to just save all the meeting details. You want to strip out some of the personal conversation out of meetings and then you want to write these more condensed, synthesized MD files. How do you go from raw transcripts that might have personal information to useful synthesized output?
Mikhail: That’s a great point as well, and that was my assumption, that you actually have to summarize and synthesize the output to be useful. But we tested this in multiple ways and it turns out that summarization actually hurts retrieval.
For two reasons. One is, when you summarize something you lose granular details, and the devil is always in the nuance and the details.
Another reason is that when you summarize, you impose a certain template onto whichever conversation that you had. For instance, the tasks that were solved, the tasks that remain, what is important, what is not important. And then on top of this template, you try to shove your transcript into this template. And we’ve noticed a huge fidelity loss. I think it was like 20 to 25% worse recall.
That’s why my personal architectural decision here was, let’s just store every single transcript, every single conversation that we had in a raw form without any summarization, because it costs nothing.
Sponsor break (27:00)
[Sponsor reads for Product Faculty and the AI employee sponsor, plus a promo for cohort 4 of Aakash’s Land PM Job program. Full details and links are in the YouTube description.]
Hybrid search and retrieval (30:34)
Aakash: Wow. And Hermes and OpenClaw on their own can figure out which meetings are relevant context and they won’t just fill up their context window with random meetings?
Mikhail: Yeah, absolutely. Because what happens is that when you ask a request, what it does is a search query. And it’s a blend, it’s a hybrid search query. It tries to do exact keyword matching. If it’s not successful at this, which I think 75% it’s not because it’s hard, it’s a very ambiguous raw context, then it does the vector search. And vector search allows you to do this exactly, fuzzy retrieval based on probabilities. And that pretty much handles the remaining 75% of use cases. So it only retrieves the relevant bits and pieces of data that are highly specific to your request, and it doesn’t really overload the prompt with tokens there.
Imperatives and the fake helpful problem (31:37)
Aakash: Very cool. So that’s the memory component. What else do people need to know to build this?
Mikhail: I think the thing after deciding on the architectural aspects of this is that it is hugely important to create the list of imperatives.
As you know, LLMs are very biased because of how they are trained. More specifically, they are focused on the resemblance of good output rather than the actual results. You can do workarounds to solve this, and the best way is to have this list of imperatives.
For instance, what for me is super critical is that there are no fabrications. If you look here, it’s just a huge list of different imperatives that we’ve created over time. It’s who you are. It’s your voice. It’s think before you act, because that’s a terrible thing I’ve noticed LLMs do a lot. They provide you the output that looks plausible but there wasn’t that much thinking behind it, because it’s full of contradictions.
Then there’s always facts over guesswork, and anti-patterns. Another angle which super pissed me off a lot, I call this fake helpful. For instance, if you ask your agent to book you a meeting and suddenly your tokens in the Google Workspace have expired, and then the agent says, well, I’m sorry I cannot do this, but it’s super easy for you to do. Just open up a calendar, type in Google Calendar, name your meeting, choose your time, choose part. I mean it’s useless. It’s obvious advice. You don’t even have to waste tokens explaining this to me. And I call this fake helpful. If you are able to put it as an imperative, it will save you a lot of time as we go.
You can clearly see we’ve made a lot of tweaks over time.
CLAUDE.md versus SOUL.md (33:45)
Aakash: So for people who don’t know, we showed CLAUDE.md and SOUL.md. What’s the difference between those files? What should be in what? SOUL.md I believe is a part of OpenClaw.
Mikhail: Yeah, it’s a part of OpenClaw. It gives effectively a specific context that is being loaded into every single prompt. And CLAUDE.md has just the highest priority of them all, and SOUL.md is the second in order of priority for OpenClaw.
Aakash: So your CLAUDE.md, it seemed like you kept that under 100 lines, which I think is the advice that Boris Cherny, creator of Claude Code, gave. The SOUL.md though is like 800 lines. So that can be longer.
Mikhail: Yeah, it’s longer. And you might argue that we are overloading the context and some of those lines might even be contradictory. But what I’ve noticed was that that’s the most robust way that gives me the most accurate output with the best recall.
And we test every single imperative. We actually test on the basic queries to see across the most important topics that we raise in the organization how good it is or how bad it is.
Tools and automations (34:53)
Aakash: All right. What else do people need to know to set this up?
Mikhail: Another angle is you need to have tools, and it has a lot of tools. Since it’s interconnected with all of your workspace, it is deeply integrated into Google, so that’s a separate tool. It’s integrated into Atlassian, that’s a separate tool. It has automations, and an automation like the one that we currently have just consumes all of the mobile reviews and writes a digest like a typical supportability team would do. And it even tags the people that are responsible and it can even raise red flags in case it’s necessary.
Claude app versus the IDEs (35:31)
Aakash: And these are all Python files. You just prompt these with natural language in the IDE to Claude and Claude writes these, or how does that work?
Mikhail: Previously I used an IDE, but now I’m doing this for showcase purposes mostly. I’m using the Claude app because it allows me to do this on the go. It is accessible from my mobile phone in case an urgent thing appears and I want to change anything within the agent.
Aakash: So do you basically create a GitHub repo that Claude Code can access?
Mikhail: Yes, exactly. So every single change is immediately committed into a GitHub repo.
Aakash: Awesome. So you do everything via GitHub, via the Claude app. That’s so powerful. People kind of have this mistaken thing that they need to use the OpenClaw gateway, but you can actually do everything through Claude and it can manage it on top of Hermes and OpenClaw.
Mikhail: Absolutely. So the only benefit an IDE gives you is that you can clearly see the structure of the project.
Aakash: And you went through two IDEs actually. So you showed us Antigravity and Cursor. Talk to us about when you should be using the Claude app versus Antigravity versus Cursor.
Mikhail: My use cases are the following. So everything that lives in the cloud and doesn’t require access to my computer, I’m interacting with it through the Claude app. But any specific use cases that I have that require my computer use, and mostly these are around browser use via different MCPs, or Excel use if I want to load a file and get some immediate feedback, or data research if I don’t really want to load it into the agent and I want to do something ad hoc, I use an IDE for this.
And why do I have two IDEs? Simply because it allows me to have two autonomous separate instances of agents running at the same time.
Hermes and auto generated skills (37:35)
Aakash: Okay, I think I understand the OpenClaw setup part of this. It’s mainly through the SOUL.md plus OpenClaw is what’s giving you that gateway to the agent in Slack. Explain the Hermes part of this. You had said it’s related to the skills and the recall.
Mikhail: The beauty of Hermes as a scaffolding is that it comes with a unique advantage over OpenClaw, which is it automatically generates skills based on the tasks that you most frequently request the system to perform. And what we’ve noticed in testing was that the recall has improved significantly. I mentioned 30%, with those automatically generated skills versus without them.
So we have it running under the hood all the time. For instance, team hiring evaluation, immigration case building, because we have different people who are relocating locally so I have a lot of questions around it. Then it made a decision that we should have a skill for this.
The same team hiring evaluation, it also made a decision. Why do I continue doing repetitive tasks? Why not just create a skill and offload all of this onto an agent? So it even takes this meta part of understanding whether or not you need a skill and makes a decision for you. That’s an amazing part of it.
How to measure recall (39:01)
Aakash: Awesome. And how did it help with the recall? You mentioned it also helped there.
Mikhail: Plus 31%. So what does it actually mean? We have five core topics, and those topics are market, business model, product key growth levers, selection, price, buyer activation, trust. Then we have more tactical things like funnel and so on.
And we have a set of different questions that we ask. We asked I think 10 questions for each of those areas and we compared the non-skill response versus the skill response. And we also looked at how accurate the response in each of those categories for each of those questions was. And what we’ve seen was that having those skills in place gave us plus 31% more accuracy, which was a deal breaker.
Aakash: How does someone set up the metrics to measure something like recall?
Mikhail: You can do it basically yourself. I’m doing this as a CPO because I think outside of the knowledge graph and percentage of context digitization, so to say, you should also be able to track more tactical metrics, which is recall.
You can ask Claude or whichever system you’re working with, or your agent directly, tell me which areas are the most frequently asked or which I most frequently interact with, and come up with 10 questions per each of those areas which are the most frequently asked. And then let’s do a comparison basically, one control group versus the treatment group with the skills, and see what’s the difference there. That’s pretty much it.
Aakash: To use sort of a derogatory word but not in a derogatory way, you’re vibing the metric. You’re in natural language. You’re explaining exactly how you want it to work and then it’s creating the metric and it can go off and measure it for you.
Mikhail: That is correct. But I mean for the first pass you have to at least visually look at the results. But once you have a methodology set up, then you can delegate the evals fully to the system.
The board skill (41:35)
Aakash: All right. Amazing. So now people know how to set this all up. Can you show us some more use cases? What are the best, most powerful things people should be using this for?
Mikhail: One cool thing, this is also around skills, but I think this would be highly practical for people especially in C level positions. So I have a board of directors that I’m in constant touch with and they make investment decisions. How much money we as a company should receive, at which point in time, what would be the payback and so on and so forth.
And there are multiple people on the board and they have different perspectives. But the beauty is, the agent was able to abstract their mental models into a set of principles, and we called it the board skill.
So in case you have a pitch deck and you want to defend the strategy for the next year, then the first thing I do is I run this pitch deck through the board skill to poke holes and give me brutal feedback on what could go wrong with the defense of this strategy.
Aakash: Fascinating. So your Granola meeting transcript is automatically recording your board meetings. Hermes has automatically created a skill around the profile of those board members. And so when you’re creating a board deck in Claude, let’s say you use Claude Design, and you would put in your input towards the end of that process, you’re going to hit it with a prompt like, use the board skill to see how the board would respond so that we can refine this. Is that right?
Mikhail: That is exactly right.
Access rights and privacy (43:18)
Aakash: Fascinating. That brings up an interesting point for me. We as CPOs, we may not want to let our PMs see what’s going on in a board meeting. How do we make sure that information is locked down so that the right groups get access to the right information?
Mikhail: That’s a great question. Two things. The first one, it all boils down to scope of ownership. The agent has different access rights depending on the person reaching out, and those access rights will define the context which this person or employee is able to retrieve, and the tools which this person would be able to work with. For instance, the board skill is available only to me and to the executive committee within the company.
And the second thing is the privacy aspect. Because not all people really appreciate that their personal meetings are transcribed or recorded. So unless you willingly provide this information to Granola, we will not train the model, we will not create the skills on top of this. So every employee has a say in this.
I’m personally a part of the experiment, that’s why I’m completely transparent and I don’t care about my privacy at all.
Aakash: Okay. So if you slip in that you’re taking care of a sick kid during the weekend, you’re okay with that hitting a transcript. But if a particular employee doesn’t want it, they can elect to take it out.
Mikhail: Absolutely. Yes. And from the get-go, we don’t really store the personal transcripts of people’s meetings, because I think that violates privacy. So unless you willingly provide us this information, we will not do this.
Aakash: So you can select, this is a product trio meeting, this obviously makes it in. This is a product review, but this is just a one-on-one between me and my designer, this does not make it in.
Mikhail: Exactly.
Autonomous prototyping and model routing (45:21)
Aakash: Got it. So the board skill is one really cool use case. What else should people be using this for?
Mikhail: Let’s see. It responded in a different language. But the thing is, we asked it, I want a specific feature, and instead it created a prototype.
Aakash: Whoa. Let’s see. Yeah, that’s quite agentic and autonomous.
Mikhail: It’s agentic and autonomous. And it sometimes gets finicky depending on which language we interact with, and we use different languages, so that’s why it might get frustrated at times and it requires an imperative. But here I think it created a prototype which is very close to our actual design system.
Which is, I mean it’s not great in terms of UX, but as a one-shot pass, I think that’s an okay thing.
Aakash: As a one shot. I mean, it’s just so much stuff it’s done.
Mikhail: Yeah, pretty much so. And it’s very compliant with the actual Nexus design system, which is the OLX design system at play.
Aakash: And do we know what underlying model it hit? Did it hit Fable or do you have intelligent model routing? How does that work?
Mikhail: The thing is, there is an agentic orchestrator which makes a decision on which model to use. So if it’s a complex request, or if it’s a specific domain area which has high sensitivity or error blast radius, most likely it will assign Fable onto it. If it’s just an execution type of work or status report writing, I think it will assign Opus 4.8 most likely, for low level work that doesn’t require super accuracy. So token optimization becomes important here.
Aakash: So mostly builds on the Anthropic stack. You’re not throwing in a Codex or a GLM yet.
Mikhail: It’s also a great point. Our engineers actually do. If we look at the typical triad, engineers have their own agentic stack, product has their own agentic stack. Of course they communicate with each other via the same RAG and MCPs.
But when it comes to product management, we use Claude Code by subscription pretty much, and we don’t really have that much token spend at the moment. That’s why it’s not a critical metric for us to track. While for engineers it is critical, and they actually have a totally different scaffolding and they have a blend of different models running under the hood.
For the breakdown of tasks, which is typically considered the work of highest complexity, how do you break down an abstract epic into tasks, they use the highest cost models like ChatGPT 5.6 or Fable. But for lowest complexity items, they use smaller models or sometimes even locally deployed open source models.
The design system built from a prompt (48:29)
Aakash: Okay. Well, how does this help you with things like your design system?
Mikhail: Pretty significantly. Let me show you Figma. So if you look here, the entire design system was built from a prompt.
Aakash: Wow.
Mikhail: And it’s very well documented and the components are robust. What is more important was that the components are exhaustive. So for instance you have all of the sizes of the components, you have all of the states, all of the relevant tokens. And it is continuously supported by the agent itself.
So in case during the conversation between engineers and a product manager there is a disconnect, because a certain component is missing or somebody tries to hardcode a component for instance, then an agent is informed about this and makes a decision to create this component in the system.
Aakash: So it’s automatically updating the design system based on the work different design teams and product teams are doing.
Mikhail: Yeah, that is correct.
Aakash: And how do you make it auto updating like that?
Mikhail: One is it has a certain tool on the product side that collects all of the requests that are coming either directly via Slack or from a coding agent. And then it just fires up a development cron job to start building those components on time.
Of course there is a review process where a designer takes over and looks through the components, whether they are correct or not, how well they are documented, are there any conflicts, are the states exhaustive and so on. But that’s a review type of job rather than an execution job.
Aakash: And is the design system skill in Hermes auto improving? Is the amount of review that the designer needs to do getting less over time?
Mikhail: It’s a great question. I’ll have an answer by our next podcast.
Backlog and roadmap management (50:43)
Aakash: Amazing. So that’s the design system. The next use case you talked to me about, which I think is fascinating and I really want to see, is around backlog and roadmap management. What do you use there?
Mikhail: You know the biggest problem with backlogs is that every single team has their own textbook. Every single team has their own Excel spreadsheet. There are many apps or tools that try to automate this but I haven’t seen a single successful case.
That’s why my understanding of a backlog is more of an abstraction on top of all of the Excel spreadsheets that teams are working with. It’s basically, every single team has their own spreadsheet. I’m not going to open them right now because of NDA things, but it’s the typical type of backlog. You have a project, you have a problem, you have a solution, different links, impact, effort, ROI and so on and so forth.
Every single team has their own backlog and they’re all interconnected. They’re all connected into the agent and an agent can fetch all of this information and can contribute directly to the backlog in case a feature is definitely worth it and passes the bar and passes the review of a responsible product manager.
The personal assistant flow (51:57)
Aakash: Fascinating. What about personal assistant? How can you use this as a personal assistant?
Mikhail: That’s my favorite part of the flow. Let’s see. I need to plug in the meeting with execs for one and a half hours next week. Not too early, not too late. Suggest something.
That’s how you would typically talk to your personal assistant. So right now it does all the retrieval processes that we talked about. And the ideal response is it will come back with a list of different time slots that would work for all of the execs who all have busy calendars who operate in different time zones all across the globe. And I would make a decision on which one to choose. If I’m too lazy, I can even delegate the decision making onto an agent.
Aakash: So while that’s running, you mentioned you aren’t even checking your email anymore. You weren’t checking your calendar anymore. So how is it managing those for you? How is it surfacing the right emails, the right calendar events, getting that prep going?
Mikhail: So every single day, I get a digest of what’s important, what has slipped in terms of Slack messages or emails. I browse through it. You can clearly see, Wednesday, July 22nd. Cleanest slot, chief. Let’s use it.
Aakash: Wow. So it’s applying intelligence. It’s figuring out when those execs are available. It’s looking across time zones and you can even schedule the meeting directly.
Mikhail: Yeah. I don’t need to open the calendar. So I’m getting the digest. I review the digest and I make a decision on which emails are actually worth acting on versus the ones that are not, or which tasks or messages I need to respond to urgently versus the ones that are not. And if I’m not feeling like responding directly, I will ask the agent to come up with a template response.
Automating the recruiting workflow (54:05)
Aakash: Amazing. I know the last area can be a bit of a headache for product leaders. On top of our own job, we need to keep the pipeline of really good PMs. How can this help with recruiting?
Mikhail: Actually quite a bit. It has three tools integrated. So the first tool is LinkedIn Recruiter. Let me try and see whether it will be able to respond. The number of product manager candidates we have in the pipeline, just the number.
LinkedIn Recruiter is responsible for reach-outs. So if you have a question, let’s say I want a specific candidate with a very specific domain expertise with this many years of experience located in this specific place, just go and find me. And if the candidate is found, then I will ask to do reach-outs using predefined templates that we’ve agreed on with an agent before.
The second thing is it is integrated into our CRM. Different companies use different CRMs. I think Greenhouse is probably one of the most popular. And I don’t really interact with the CRM anymore because I do all of the work via the agent. I just type in, who is in the pipeline, or I provide the feedback directly to an agent. An agent posts this feedback or moves the candidate along the funnel. Or if needed, I can even build the funnel statistics to see where the gaps are in terms of the hiring.
And the third aspect to it, which is quite amazing, is the interview process. Because all of the interviews are transcribed, you can get an immediate third person opinion on whether or not the candidate really lived up to your expectations. And in case not, of course you are the final decision maker and not the AI.
Aakash: Yeah, it just gives you an alternative point of view.
Mikhail: Which is sobering in many cases. Then you can ask an agent to draft a rejection letter and reach out to the candidate. So this automates probably like 70 to 75% of the recruiting workflow. And more importantly, it drafts beautiful tailored rejection emails with really deep focus areas for improvement that are derived from the Granola transcript. None of the recruiters that I know actually do this well.
Aakash: So you deliver a much better candidate experience as well as making it faster and easier for you, and potentially even making better decisions because you have the transcript.
Mikhail: Exactly. Yes.
Will product management exist in the future (57:13)
Aakash: This has been mind blowing. I want to talk to you about some hot topics now.
Mikhail: Sure. Let’s dive into them.
Aakash: What’s on everybody’s mind right now just in the zeitgeist is what’s going to happen to my job. You just showed me how you might be able to do the work of two PMs with one PM. So what’s going to happen to the future? Is product management going to exist in the future? What’s your take?
Mikhail: I have good news for you. I think not only is product management going to exist, I think it’s going to thrive.
Because product management in large organizations became burdened with layers of hierarchy, with layers of corporate processes. And in many instances product management is about compliance, performance, status reporting. So it’s suboptimal product management, to be honest.
And I think what AI really helps you to do, it doesn’t remove the decision making from the process, it gives you the juiciest bit. It removes a lot of the operational overhead from your work. And what product managers are now more focused on is value discovery. Because what AI doesn’t really know is what your customers want, because AI doesn’t really understand the customer journey well. It has the funnel metrics obviously, but it cannot really go offline and talk to a customer. You have to do this.
And that’s the funnest part of the job, because you don’t do all of this theory anymore. You’re only focused on delivering the real value. And I don’t think it’s going to go away anytime soon. What actually will happen is that the teams will be smaller, leaner and faster.
Aakash: So do you anticipate that the ratio of PM to engineer is going to change from what it historically was?
Mikhail: Yes, definitely. I think not only is it going to change, the lines are going to blur between what a product manager does and an engineering team or an engineering manager does. There’s likelihood that they could overstep and be a support system. Because at the end of the day, if a product manager is an orchestrator of the product quality and token budget, an engineer is an orchestrator of execution quality and token budget. The same applies for the designer. A designer is responsible for the consistency of a design system, which means quality and token budget as well. So the boundaries between those roles would be very blurry.
Staffing a new product area (1:00:01)
Aakash: So you’re a CPO, you’re thinking about maybe I need to create a new product team in a particular area. When I was a VP of product at Apollo, we would typically think about, okay, if we’re going to create a new front-end product area, we probably want five to seven engineers. We want a PM. We want a designer. When you’re thinking about a new product area now, how do you think about staffing that?
Mikhail: It’s a great question. So I typically divide them into two different groups. The first group is high complexity and high error blast radius. I think for those specific domains nothing has pretty much changed, and it’s all around monetization. It’s all around algorithms like search, ML, where a single tweak might have a dramatic impact on the conversion or on the outcome. Here you would need to have a dedicated owner.
But for the remainder of the domains, like customer facing domains, what I’m currently seeing is that teams can be easily scaled without adding people. You might have one product manager owning three or four domains across many platforms at the same time.
How to hire AI native PMs (1:01:20)
Aakash: Okay. So the PMs you’re hiring, they feel to me like they’re truly AI native PMs. Even if they’re not building AI features, you’re probably hiring people who are really AI forward. What’s your process these days to hire PMs? How do you find PMs that are at this level of AI native, that they can actually succeed in this environment?
Mikhail: Great point. So I typically look at three things. I think the fundamentals haven’t changed, but I also added additional qualifiers into the interview. The fundamentals meaning problem solving and systematic thinking. If you could have both, you will adapt in any situation.
But the additional qualifier that I’ve added is a level of craft when it comes to AI. How deep a product manager is in AI. And there’s a very simple way to test for this. I’m asking, which of the use cases of your daily work have you automated using AI. And I get a spectrum of responses.
If your response is on the level of, well, I’m talking with ChatGPT using a terminal and using a web interface, to me that’s a low level of immersion or maturity when it comes to AI tool understanding. But on the other side you might have a very advanced AI agent orchestrator, a product manager who has automated all of his professional work life and built it entirely using AI and so on. And then I can drill deeper, asking how do you evaluate this, how do you make decisions on tools, on the brain, on the memory and scaffolding and so on.
But I’ve got to be honest with you, in the market where I operate, it’s a rare skill set. So product managers who are able to effectively answer those questions, they get way more points.
Where to find Mikhail (1:03:28)
Aakash: You guys heard it from him, not from me. He is a CPO at one of the greatest companies I would say to work at and he is running it in an AI native way, and this is the skill set you need, the skill set we taught today. So if you haven’t already, go play around with Hermes, go play around with OpenClaw. And if you’re a CPO, steal Mikhail’s playbook. His team is operating at a level unlike many others.
If they want to learn more about you, get in touch with you, find your content, where can they go?
Mikhail: Well, I do have a Substack newsletter. It’s called Corporate Waters. You can reach out and get immersed in my thinking and the actual use cases. Outside of this, you can reach out to me directly over LinkedIn.
Aakash: I highly recommend Corporate Waters. If you guys didn’t know, Mikhail and I did a really fun deep dive last year where he looked through all of the interviews he’s done in his lengthy career, which by the way, before he was a GPM at Bolt, he was also an IC PM for over a decade at companies like Yandex. So he collected all that data and we looked into what interviewers are looking for. So if you want to go see more information from me, you can find it in my newsletter or his newsletter.
Thank you so much for sharing all this amazing sauce today.
Mikhail: It was a pleasure, Aakash. Thank you for having me.
Aakash: The craziest thing is that Mikhail agreed to open source the information that he used to build all of this. So go check the link in the description for the GitHub repo on his profile. You can fork that and you can begin building this company operating system for yourself.