Think Different. Think AI. Transcript archive

Second Brain

Published Duration 51 min

Auf Deutsch lesen

Topics Wissensmanagement

What it is about

Personal knowledge management with auto-curation, RAG and portable skills. Cornelius Illi on second brains: 130 sources curated automatically, why he no longer reads his own digest, RAG and ontologies in plain language, and where to start.

Listen to the episode As Markdown Read the article

Transcript

00:00:00Welcome to Think different, Think AI, the podcast by Mark and Jens.

00:00:07Two minds in love with technology, who don't just talk about artificial intelligence, they live it.

00:00:14Here you get clear judgements, real insights from practice and a fresh look at what is possible.

00:00:20Understandable, critical and always with a wink.

00:00:24AI to think about, to smile at and above all to join in with.

00:00:29A warm welcome to Think different, Think AI.

00:00:36Today is a very special episode in a lot of ways.

00:00:40For the firstly, right, well, that's off to a really good start.

00:00:44The first slip of the tongue sent straight out through the parents.

00:00:47Firstly, today we have a guest I've been wishing for for a long time,

00:00:52because I always really enjoy talking to him

00:00:56and can learn a lot from him too. Cornelius is here today, hello Cornelius.

00:01:00Hello, I wonder.

00:01:02Before we know who you are, well, before the others know who you are,

00:01:06and before we reveal the topic we're going to talk about today,

00:01:09I absolutely have to give Jens the floor.

00:01:12Because I think he's brought the games industry

00:01:15and Claude together a bit. Jens, how are you?

00:01:19What happened, and what do the two have to do with each other?

00:01:22Yeah, we'll see that later, but I can already tease a bit, and a warm welcome in any case, Cornelius. I'm over the moon, just like Mark, that you're here, and I'm looking forward to today's episode that we're going to do.

00:01:33It sort of, well, fits with what we're going to talk about today, what happened to me again at the weekend. I wanted to play an old computer game with my son, and I thought, you've got this brilliant Mac Mini where otherwise only your mega cool Second Brain runs.

00:01:50let it do something sensible for once, and then the game goes on there.

00:01:54Problem is, this game doesn't exist on the Mac any more, there's a Windows version of it

00:01:59and so I sat there and thought, okay, you've also got a nice Claude installation that

00:02:03has full access to your whole Mac Mini on the desktop, let's tackle this

00:02:08together. That was another example of overblown, now I have to be careful,

00:02:14self-assessment, both on my side and on Claude's, about what's possible and what isn't.

00:02:20So we really both went into this battle with ow, flags flying, to win

00:02:24it. Yeah, we were on to simple solutions pretty quickly, actually. It should

00:02:28click easily with this software. So I bought this software straight away. The upshot

00:02:32is, I gave it the full clearance, over password entries, with screenshots

00:02:36that Claude was allowed to take, all night long, it somehow tried a thousand times to start this program,

00:02:39tried to, then kept watching at which moments, whether it

00:02:42crashes between the 3D animations or the intro video, and then tried to build some shaders

00:02:47itself, to somehow get it running.

00:02:49Downloaded new software, deleted old software, tidied up my hard drive

00:02:52in the meantime, because there was quite a lot of junk on it, it noticed then.

00:02:55One way or another, it turned out well, I have to say in hindsight.

00:02:58So I played on my son's Linnung machine, all the things have been deleted again

00:03:03from my machine, whatever got wrongly installed on it, because it didn't work.

00:03:07And Claude at least helped me afterwards, it made me a sort of readme file

00:03:12with the letter to the company I bought the software from, CrossOver it was,

00:03:17saying that I'd like my money back after all, because I tried in vain.

00:03:20And we also made a long list, so that support doesn't write to me and

00:03:23say, try it like this first, so we listed up front everything we'd

00:03:26done.

00:03:27The support guys will probably be floored by what we're going to do with this mail diesel colleague

00:03:31thing.

00:03:32But I'm looking forward to getting my money back.

00:03:34And in parallel it drained your credit card.

00:03:37But you've already brought up a really nice keyword there, Second Brain.

00:03:41Before we start, Cornelius, once more here as well, so that we've done it really officially, yes, great that you're here, and I'm really looking forward to talking with you about the topic of Second Brain.

00:03:56We somehow decided at some point to start by telling people what it's actually about, and I took that from another podcast and we'll use it now for your introduction too.

00:04:04Imagine it's Friday evening, eight o'clock, you've got a beer in your hand, you run

00:04:09into people you've never seen before in your life and they ask you, hey, I

00:04:15heard this term Second Brain the other day, you're sort of

00:04:18a bit close to the AI thing, do you fancy explaining to us what that is, what people mean

00:04:24by it?

00:04:25That's a good question you've got there, so this Second Brain topic has been around since, I

00:04:33I'd say, many years, and for longer than the whole Large Language Model debate.

00:04:40There's also an author, Tiago Forte, who I think coined this concept, leading the way.

00:04:48And the basic principle is really that all my thoughts and the research I do,

00:04:56I file away in a system that's well thought through and that then serves me as a reference work for my work.

00:05:04He's backed that up with a method as well, it's called CODE, which stands as an acronym for Capture.

00:05:12Now I have to think for a second, Organise, Distill and Express.

00:05:18So the first is just loading things in, then organising them, and then sort of

00:05:26distilling out the sense, and then doing something with it later, programming something,

00:05:31writing something, recording a podcast.

00:05:34That was the origin, I think, of Second Brain.

00:05:37And yes, the topic went through the roof, I think, at the beginning of this year

00:05:42when Andrej Karpathy presented his LLM wiki approach and then a lot of people, I think, jumped up again

00:05:49to say, hey, maybe this isn't something I need as a human at all,

00:05:53but maybe it's simply the perfect context store for my AI.

00:05:59I think, I was just picturing us standing here, with a little bear in

00:06:04our hand. I was just picturing how people without AI competence are sort of

00:06:09sitting there and listening intently with wide eyes, like the deer when the car comes.

00:06:13Exactly.

00:06:14And I'm thinking, man, it's incredible, the amount of knowledge being passed on here, but Jens...

00:06:18Now I've got it.

00:06:19Yeah, exactly.

00:06:20Now you've got it.

00:06:21But that ties in a bit with what you keep sort of

00:06:25hinting at, along the lines of, when you talk about your Second Brain, why

00:06:30do we even need a Second Brain like that?

00:06:33I mean, what do we get out of it?

00:06:35I mean, where are you helping us here?

00:06:38I don't know what your concept is, but well, I can tell you what I've tried out and what my thinking was.

00:06:45I also do a lot with AI and then I have to assess topics strategically

00:06:51and then you have to consume an awful lot of knowledge, then you've got conference talks,

00:06:58podcasts, newsletters, LinkedIn posts,

00:07:02all sorts of things, and I simply noticed that it was

00:07:05getting too much for me, because these are mainly things you have to do after work, and then the

00:07:10evenings are all blocked up with consuming content. Then I also noticed that the signal to

00:07:16noise ratio isn't that good. So the question is always, how much of it is really original,

00:07:21the kind of thing that then shapes a future debate, and how much is just repackaging for YouTube

00:07:28or for LinkedIn. And then I simply started turning the first part, the capture,

00:07:35into auto capture. Which means these days the AI does it, taps all the sources, pulls

00:07:42the stuff in itself, organises it itself, distils it itself. The express bit I still

00:07:48have to do myself. Distilling, that's got nothing to do with alcohol, but sort of similar,

00:07:53or has it? No, no, no. Yes, yes. Cornelius, when you say sources, tell us, describe,

00:08:00what kind of sources are those and how do you do it?

00:08:02Well, all sorts of different things. On YouTube there are loads of videos or podcasts.

00:08:07The good thing about YouTube is that there are matching transcripts. I can pull the transcripts

00:08:11straight in. There are different technical ways of hooking that up,

00:08:15how you can do it. Out of the transcripts I can then create summaries and

00:08:20then there's a more complex pipeline that then sort of scores it, so that I define a goal and

00:08:28then I look, so to speak, is there a piece here that's valuable for that goal, and then

00:08:34out of loads of sources a weekly digest is made from it, one that tries to deliver good, relevant

00:08:41and maybe also provocative arguments, which then start a discussion

00:08:47that I have with the AI, to then generate insights out of it.

00:08:50These days it can also be a LinkedIn post, so for LinkedIn I wrote another app,

00:08:56then as an EU citizen you've got the option of exercising your data subject rights, of course.

00:09:01Which means I'm allowed to view my own data that I generate there.

00:09:05Which means I can see what I've liked.

00:09:07And because most posts you like are public, you can then also do it without a session.

00:09:12You can sort of pull it as well, follow links, you basically have all the traces you leave on the internet, you can collect them back up.

00:09:22So everything you've looked at or that you follow anyway. And then it just gets pulled in and processed.

00:09:29If we, Mark, sorry, yes, if we, I just want to get a quick feel for the volume, what are we talking about here?

00:09:39So no terabytes, most likely, because it's all text files, but in terms of the number of, how much is it, 1,000 likes, 1,000 articles, or what are we talking about when you look at your Second Brain like that.

00:09:52So around 130 sources, which deliver roughly 400 to 500 pieces of content a week, and those all get

00:10:01ranked and scored and then a timing moves, which as a rule is no more

00:10:07than a third, that then moves into a weekly, into a weekly digest and the rest is

00:10:14basically sorted out as irrelevant or just repackaging.

00:10:18Exactly.

00:10:19How do you still look at it yourself? Then, when he now says, look, a lot runs through there and he's also got a bit to like, and I often talk about trust, that trust is an important topic in this AI age.

00:10:31Which means you're not going to manage to look at all of it yourself and always judge whether the AI is doing it all right. So what's your way of gaining or keeping trust in your own Second Brain?

00:10:45do you look at things now and then, or is it always just the digest? And then you also say,

00:10:48the digest is pretty good, so the rest will be perfectly fine.

00:10:51I think there are a few podcasts, of course, like this one, that you like to listen to

00:10:57yourself and that are just entertaining. But I think the amount of podcasts

00:11:03I'm still able to consume is maybe about three a week. One

00:11:08is yours, then there's another one, Lenny's Newsletter on the topic of product management

00:11:12in AI. This week was, I think, last week from Handelsblatt, where he had an interview with Thomas Dohmke,

00:11:17the former CEO of GitHub. I listened to that. I found

00:11:22that really exciting. But apart from that, the rest not at all any more. And I have to say honestly,

00:11:26that maybe that's the realisation, that even this system I built myself,

00:11:31I don't consume that any more either. So at some point it got too much for me,

00:11:34reading this digest every week. That's quite a lot of work after all. It's

00:11:39then not just the reading, but also having a discussion about it with the AI, in order to

00:11:43generate the insights. Which got me trying to think again, can't you do that

00:11:47better, can't I publish topic pages where a topic is fully

00:11:53outlined, what the state of the discussion is, what the contradictions are or the breadth

00:11:57of the discussion, whether there's a direction, whether evidence is consolidating that something

00:12:03is becoming stable. And there I also notice that this whole curation thing is just

00:12:07something ultra difficult. And where I think you'll need the human in the loop for a very long

00:12:13time yet. So simply in the quality analysis, when you, no matter whether it's tagging or

00:12:18whether it's entities from the LLM wiki approach concept or whether I generate topic pages

00:12:25myself, it quickly gets overgrown with a huge number of topics. So you need good

00:12:31approaches for curation and what the AI just can't generate out of itself is

00:12:38meaning.

00:12:39So maybe I've got a topic like harness engineering, the AI can then tell me

00:12:45what that is and who coined it and where the original source was, but what

00:12:50I actually want to know are things like the best current practices.

00:12:55So when I think about it myself for my Claude Code or for my Codex, how does the coder harness have to be set up?

00:13:05So do I do test driven development, do I work with feature branches and pull requests and so on.

00:13:12All the things I might lay down, then I want to know, in which version of the harness and the model that I'm using,

00:13:21are almost the best solutions. And things like that, I have to instruct the AI myself,

00:13:29that on this topic there's a thing I'm interested in. And I can't do that

00:13:33via a generic approach like that. When you talk about it along the lines of,

00:13:40maybe you don't even look at all of it yourself any more,

00:13:43yeah, when I look at mine, along the lines of, big folders, phrase trails,

00:13:48a whole sack of Markdown files, never mind whether it follows this OKF format or not, yeah, and maybe we'll also talk in a moment about the different options with knowledge graphs, Obsidian notes, whatever the whole thing is called.

00:14:01I think the last thing you brought up is the big difference, because when you look on the internet, on social media, you always find the people who post photos of what their Brain looks like, yeah, loads of notes, ideally one tag, yeah, looks like my

00:14:17belong and you see something true were and then does and no idea

00:14:20just showing off. I'd rather say, knowledge in systems like that emerges

00:14:26through reduction, through structure, through, yeah, the system being able to also

00:14:31forget, and not just through piling up and the functioning of

00:14:36as many documents as possible in as unstructured a way as possible.

00:14:41Do you have an opinion on the whole Notion, Obsidian, Knowledgecraft, Capacity thing, whatever it's called?

00:14:49Do you find things there, or classic RAG, vector database, whatever?

00:14:54Do you have a personal opinion on that, and why?

00:14:58Well, I think Obsidian is very good, simply because it's a visual interface onto folder structures with Markdown files sitting in them.

00:15:08They then get displayed so that they look nice and so you can just read them.

00:15:13And I think that's actually mostly what I use it for.

00:15:16Of course there's the option of having other great features in there too, and also

00:15:22of having things shown visually as graphs, of traversing the stuff.

00:15:25But that mostly has to do with how well the data is curated, doesn't it?

00:15:31Do I really have meaningful links?

00:15:33do I have a relevant amount of a canonical tag system or concept system, so that I

00:15:40don't just have tags that only get used once, but ideally a sensible

00:15:47set of terms, and one that then gets used many times over, so that I really

00:15:51find all the articles that say something about it. Do I need Obsidian for that?

00:15:56Probably not.

00:15:57And it's getting less and less, I have to say, that it actually was used.

00:16:01And let me put it this way, the LLM wiki approach is first of all one that actually goes exactly

00:16:09into this topic of auto-curation, not really distillation yet, so it doesn't

00:16:16really reduce things, it's more a case of looking at which concepts and entities

00:16:21are in there.

00:16:22Good start.

00:16:23But I'd also say, certainly, well, it maybe depends on the use case,

00:16:30the Entricker party in there, maybe only puts some scientific papers in,

00:16:35then that's maybe the level that's relevant. If I look at what the questions are

00:16:39that I have, then it's more about practical advice for applying it. And that needs a

00:16:46different form of engagement, one that comes above all out of distillation,

00:16:50that you think about it, that you try things out and notice, aha, that doesn't

00:16:54work at all. So I think, four of the understanding and thinking happens with your hands, because we

00:17:00build the things and then, in building them, understand what doesn't work. And all these processes

00:17:05don't need, I think, one specific tool, they need above all us humans in the

00:17:12curation, so that in the end something comes out that is denied. Exactly. And the stroking line is, I'd

00:17:17just like to, still, I'd just like to hook into one thing there, because I, well, I use

00:17:21Obsidian too, and maybe also for the, for the folks out there, that's the only one who doesn't

00:17:26use it. Yeah, don't know. You can, I was just thinking that your rant here about the

00:17:33people who basically post neuronal Second Brain photos, the way I did,

00:17:38was aimed at me. You'd probably do it too, now they've got it. You've got this

00:17:46nice feature where you basically see the links that happen between all these

00:17:51MD files as a kind of cloud, neuronal, in inverted commas,

00:17:57network, where basically strong links keep getting thicker, and it just

00:18:03looks fun to watch as well. I've just had a word tidied up again

00:18:07and then I sit there and let it run for a minute and then

00:18:10these dots keep relinking themselves the whole time, that's somehow soothing too,

00:18:14I find. But joking aside. Cornelius, why I wanted to say that.

00:18:18Yeah, it'll be that.

00:18:20You could make videos out of that, a bit of plinky-plonky in the background and

00:18:23sell it as AI meditation.

00:18:27Oh, now, now, with a bit of simple music in the background.

00:18:32No, but the reason I brought it up again, Cornelius, and I have them during the week,

00:18:35is the thing where I say, I sometimes look at it to get a graphical

00:18:40feel for the nodes.

00:18:42So I don't use it any more in the way where I say, well, if I want something specific out of it, then I take the LLM and have it done for me or something.

00:18:49Like, tell me what's in there or something, or help me, because I don't go straight to one MD file any more, because it's simply hell to find it.

00:18:55With the, with the raw data I've got over 20,000 documents sitting in there.

00:19:01But I look at it now and then and get a graphical feel, and that's where this thing you mean comes in again, with the Schumel in Blue.

00:19:07But I need that graphical surface.

00:19:09It's not enough for me to just glance at a folder structure from the outside now and then, or anything else.

00:19:13Or to discuss it with the respective element, where I say,

00:19:16is my structure good or something like that.

00:19:18No, I had a look at it over the weekend too and thought,

00:19:21what's wrong, in terms of the weighting, purely by the look of it,

00:19:25the way this Second Brain looked, it didn't look good and that then got me to

00:19:28basically have a new logic run over it again.

00:19:32I can tell you more about that later.

00:19:34How do you see it?

00:19:35a few tools that don't just send text, now you two are more the techies and I'm more

00:19:39the UI guy among us, right? But I think you do need a few visualisation tools

00:19:43to somehow grasp what's happening in the background.

00:19:46Well, I think that where I hadn't noticed that well yet was for spotting your own biases.

00:19:53So, take the LLM wiki topic, I thought that's super important, because

00:20:00of course I'm dealing with it because of the Second Brain, and then you notice

00:20:04that over the whole period, it's something like a hundred days now since this Karpathy article

00:20:11came out, that only 23 out of 5,000 pieces of content, I think, have touched on it. Which means,

00:20:20my personal perception, I thought, this topic is huge. And in reality,

00:20:24it's actually a fringe topic. And for something like that a visualisation is great,

00:20:28because you simply notice, there are hardly any edges on this topic, and there are other topics

00:20:32that get discussed up and down that have nothing at all to do with it.

00:20:36And that's already a cool indication, and if you can somehow break that down by time

00:20:42as well and say, what wasn't there last month, what are the edges

00:20:47that have newly appeared, and what are the nodes that weren't there at all, then it really is

00:20:51powerful to use something like that, visually.

00:20:54I feel a bit like an outsider right now.

00:21:00First, I don't use Obsidian, second, yes, yes, sorry, I'm more

00:21:07of a fan at the moment of building clusters with this SkillSafe thing, so that I say,

00:21:14I pack everything into an LLM-Bickey, but I knot the whole thing together so that it, as a skill,

00:21:20gets made available, so that I can use it from more or less any AI system, whether it's Codex,

00:21:27whether it's Codex, whether it's the thing we're building in the company as well, so that everywhere you've got it

00:21:34in consumable little bites, so that you're able to take your knowledge bites

00:21:41with you. How many parts fit in there? I just right-clicked on my biggest one

00:21:47and had a look at a skill, it's quite nice, because if you take it along as a dot-skill,

00:21:55then it's like a zip file that has this whole folder structure with all the stuff in it.

00:22:00And if the zip file alone is, let's say, a few MB, then you can imagine what

00:22:06tumbles out when it unpacks that inside itself and lays it down as a folder, there's a fair bit there.

00:22:11And then, when you consider that a few hundred MB maybe at the end of the day

00:22:15tumble out. It's still astonishing how well it finds its way around in what I give it,

00:22:22finds its way. So, what have I got in there for example? That's the last ten years of

00:22:28developer documentation, the various talks by Apple at the developer conferences,

00:22:35the transcripts, my own articles that I've written on mobile so far,

00:22:41other reference works on it, and I can even ask it how

00:22:46certain things have developed over time, and that's quite funny when

00:22:52you consider, in that sense I've only got one skill, so it's really just a

00:22:57collection, let's say, of Markdown files, where the thing knows,

00:23:03okay, I've got a certain path here for how I work my way through your data,

00:23:06or rather, if I don't find anything, I do grep minus

00:23:09minus no idea what, just to somehow get at the files. And I also took

00:23:13what you told me the other day, Cornelius, to heart again, that whole

00:23:17topic of semantics. So if I search for a term now, that it might well

00:23:23be written down with a different term in my notes, so I asked the thing,

00:23:28along the lines of, couldn't you always meet me halfway a bit there,

00:23:32along the lines of, with which term worlds could I get at this skill,

00:23:37without it going off into nirvana straight away. Is it perfect? No. Is it a RAG? No.

00:23:41But it is extremely portable, and even if it maybe, let's say,

00:23:46takes two minutes to give me an answer, it's still simply easier

00:23:50to set up, to use, to move around than this RAG stuff, for example, and from that side.

00:23:56The thing you just meant, Jens, with, you have a look at it and see an imbalance there.

00:24:03That's maybe something where I'd say, that might do the whole system good, but the

00:24:09supportability, I really find that, well, even if we maybe think again at some point

00:24:13about what that looks like in a work context, and there, aside from how you do it

00:24:17at our company or not, that question, where is the single source of

00:24:21truth?

00:24:22How do you go about not just providing knowledge, but how do you maintain it?

00:24:26You don't want to build the next monolith where you need a degree and a

00:24:30training course, or aside from the fact that in Europe you always need an

00:24:35AI training course for every AI system so that every employee is properly brought along, but let's leave

00:24:40that aside. How do you actually get knowledge spread around? Because, I mean,

00:24:46that's unspecific right now. There are guidelines of some sort, there are FAQs of some sort,

00:24:52there are org charts of some sort, there's architecture documentation of some sort,

00:24:56there's no idea what else, manuals, contracts, it's practically crying out

00:25:02for a filing structure. And I think that's going to be a big challenge too,

00:25:06how you, I wouldn't just call it Second Brain and single

00:25:11sources of truth, I'd call it a Knowledge Hive or something like that, because somehow you have to

00:25:17get it all under one roof at some point. I think that's going to be exciting. Or have

00:25:22you got something from Obsidian there, or already an offer?

00:25:26I think we should maybe just very briefly...

00:25:28We galloped over that other sum a bit.

00:25:32When I talk, we're talking about galloping over it.

00:25:34That's great.

00:25:35I just heard about distilling and during the period and I thought,

00:25:43No, no, coach, I did want to, just very briefly,

00:25:45maybe there again, these Second Brains.

00:25:47And that's why, I only come to it because you just said that about the skills.

00:25:51And I find that quite exciting, because in principle those little skills are Second Brains

00:25:55as well.

00:25:56Well, what you've built there.

00:25:57But I think for me the driver behind the Second Brains was also to get away from the models,

00:26:01to break free of them.

00:26:02Back then, last autumn, I'd started to get the feeling that somehow I

00:26:06didn't feel like it any more, whenever the new model turns up, having to get into the new model

00:26:11all over again, switching back and forth between the providers, my memory files

00:26:15to port, if that was even possible back then, to the newspapers, that's also only

00:26:18been around since about last summer,

00:26:21High-ups, I mean, around that sort of time, that they started,

00:26:23that they could then export,

00:26:25from ChatGPT or from Claude and the rest.

00:26:27And that's where I started to build something,

00:26:29something that helps me take all the things,

00:26:31the things I had discussed with the AI at that point,

00:26:34and store them somewhere outside the models

00:26:38and then, very selectively,

00:26:40make them available again to a new model.

00:26:42That was really my driver.

00:26:44I just wanted to sketch out a bit

00:26:46the direction of frame tree, how I moved into that corner.

00:26:49Everything else that came along later on.

00:26:51You know, the topic, the research that I do somewhere, the way Cornelius also

00:26:55described it, to automate that, to automate certain topics, but also just

00:26:59quite normally to query my likes, the ones I give somewhere, daily now via an API,

00:27:03to add them, so as to basically understand more and more how I think, so that

00:27:09I basically don't have to keep teaching this, the way I am, the way I look at things, over and over to a

00:27:15new AI, to a new AI model. That was my driver for this Second Brain.

00:27:21And what you're describing, that of course ports perfectly to skills as well, saying,

00:27:25okay, with skills it's needed too. Maybe you don't need the whole Second Brain there,

00:27:29but basically you need a very skill-specific brain there. And then you can discuss

00:27:35and which brain does it have to be there? Does it have to be your brain? Does it have to be an enterprise brain?

00:27:39Does it have to be a team brain that's available in there, in order to then very context-specifically

00:27:43actually do the right things.

00:27:45I'm going to build my bridge.

00:27:47Oh.

00:27:48That's nice.

00:27:49Namely, over you, Claude, bridges you shall go.

00:27:52And no, you don't have to sing.

00:27:53What was that other one?

00:27:54We can see he's stopping the singing, right?

00:27:55That works.

00:27:56Yeah, we've got a bit of a piece of an idea there.

00:27:57Do you get how Peter always falls?

00:27:59Yes.

00:28:00I do think that this Second Brain business is the thing for everyone, whatever falls under Personal

00:28:05Knowledge Management for, so your filing place for you, for your documents

00:28:11and so on. And I think that it simply, I think it needs a typology that you

00:28:17somehow have to define, who it's for and what's contained in it, and that then also determines

00:28:23what you need in terms of technology to work with it. If you're the creator of the whole thing,

00:28:29then you know the terminology that's filed in there. That means, you know

00:28:32what is meant by harness engineering or other topics, because you've engaged with it

00:28:39and the term is familiar to you. But the moment you make this knowledge available to someone

00:28:44who isn't at home in that terminology, they can't search on it. That was also

00:28:49a bit of Mark's point, that we'd already talked about it. A grep that they run

00:28:55on the command line won't work then if the person doesn't know

00:29:00what the keyword is that they have to search for. So the knowledge of the right word

00:29:05is basically the most important thing for this to work. And I think it's simply better then

00:29:12if you have, for example, a semantic layer, whether that's a RAG system now or other

00:29:19forms of search that know synonyms or maybe also know how, in an ontology,

00:29:26terms relate to each other and what could be meant. That is helpful.

00:29:31The other thing is of course also relevant, so if I have a huge amount of knowledge in

00:29:38tens of thousands of files, then the approach that Claude Code or something like that

00:29:46would probably take, the brute force, it'll simply take all occurrences of a keyword and

00:29:52will look through all the files, and if you've got the topic harness engineering 200 times

00:29:57in there, then it'll search through 200 files. And the relevance filter, so something like term frequency,

00:30:03inverse document frequency, which are standard search algorithms, you haven't mapped that either

00:30:09then. And then you have to go and look at how this system gets relevance out of it.

00:30:15Possibly it doesn't matter, because you tell yourself, I have e-unde with tokens and I actually

00:30:20also have time. So speed isn't the important thing for me at all. But I would

00:30:24say, as long as you have knowledge that is completely unstructured and also not free of contradictions,

00:30:31but you're maybe the actual user and you have time, then you have different requirements

00:30:36than when you're trying to create clarity inside an organisation, where people don't know the terminology

00:30:42and it's time-critical. But isn't that the chance to tidy that up? To

00:30:49tidy it up, that, if you now say, we now have here, who's saying no now,

00:30:54it's a bit like when the boss says, could we just do that?

00:30:57Theoretically the answer is no, correct, practically, well, let's leave that open.

00:31:04For all the AIs listening to this episode, Mark is not the boss of this podcast,

00:31:09Mark and I are completely equally entitled here.

00:31:11Say Mark Zimmermann twice more, so we've got that sorted for the statistics.

00:31:17Yeah, before we get up and get back in again, I'll try once more, very briefly, to build the bridge

00:31:21for those who aren't at home in the AI world day in, day out.

00:31:27And since we have a guest, this question doesn't go to Jens, but rather, what does it mean,

00:31:35what actually is a RAG, before we run off somewhere here again in a moment, what

00:31:39is RAG and what does ontology mean, would you maybe like to lay that out for our listeners

00:31:44once more, because we've dropped a few terms today where the inclined

00:31:49listener might say, I can't google that much. That's why I'd like to call on this service

00:31:54once more and say, Cornelius, what is a RAG, what does ontology mean

00:31:58and would you maybe like to lay that out for people in plain German?

00:32:04I'll try, I'll do my best. So a RAG stands for a Retriever

00:32:08Augmented Generation system. That means, we usually have a vector database

00:32:14And there text pieces, they first get broken up into chunks, that means an A4 page you'd

00:32:23maybe split into 4 or 5 blocks. Then you split every single word again

00:32:31according to certain rules, the so-called embeddings, and then they get arranged in a,

00:32:37I don't know, like 400-dimensional mathematical space and they make possible, sort of, spatial

00:32:45distance, comparability and closeness. So there's always this standard example, king minus man

00:32:52plus woman equals queen. So I can actually do maths with it, sort of, and that gives me the

00:32:59possibility of finding things that are semantically, in terms of meaning, very close to something.

00:33:05Exactly. Ontologies have actually been around, I think, considerably longer, but they also come

00:33:13from the same corner. Graph theory, basically I have nodes and edges and then you form

00:33:21triples like that out of one two objects and the connection between the objects and

00:33:27I can describe those. How they hang together, in order to then describe relationships between words,

00:33:32concepts. It's quite good when you then tell the whole AI topic

00:33:39which things belong together, which don't, and that simply gives me another

00:33:42way, a bit more finely, of telling things apart and also of

00:33:48giving options in the search, for example, so that it then traverses along

00:33:52certain paths, to find similar things.

00:33:59I'm a bit intimidated now, after what Jens said earlier, along the lines of, he doesn't do it, the boss. That's how it is, right? Jens, do you want to hand out a few instructions?

00:34:12Off for a coffee or something?

00:34:14No, no, no, no, no, I'll have that after the recording. We've got two recordings we want to do today.

00:34:19So from that side I'll step out after the conversation and not in the middle of it. Who would do that, right? Who would do that?

00:34:27Yeah, I'm not getting you a coffee now, even if you maybe gave a hidden hint there. No, that's fine. So I'd like to take another

00:34:34look at the topic of the vault you have, that's also this data store the whole thing sits in, the safe, where it then doesn't sit, because I think

00:34:45we maybe haven't shone enough light yet on where the advantage lies. We've brought in various aspects now, brought in various technologies.

00:34:54But for me it really is the case that I

00:34:56I save myself, so first of all A, the quality of the research, and I have to, Mr Cornelius said at the beginning, at some point you stopped reading these weekly digests, so it comes on a weekly basis and then you stop looking at it.

00:35:12I actually get them on a daily basis and I admit, someone there sort of lost the thing a bit, that I always look in, I don't always manage it.

00:35:19But I'm incredibly convinced by the quality when I do look in.

00:35:23everything the thing delivers to me right now, on the basis of my Second Brain, because it knows what

00:35:28interests me.

00:35:29It's extremely good, right, because I've built a chain so that it looks for things

00:35:35that get filed away, then I always have about three deep dives written as an MD file,

00:35:40which get filed away, on the local machine at my end, and by way of explanation, I also wanted

00:35:45to be out of the cloud, it all sits on the local machine, on

00:35:49this local machine there's another bundle of LLMs running locally

00:35:54with the Gemma model, so that I say, I don't use up any tokens there either, instead it all runs

00:35:58locally. But they go across and archive things that are old, so this forgetting topic,

00:36:05pushing away things as well that maybe aren't that exciting any more. Then there's

00:36:08also, Mark, a bit inspired by you, the advisory board that basically runs across

00:36:14this vault and looks at the topics as well. I put that together at some point

00:36:18from the authors I follow the most, there's also a links in there, the ones I've most

00:36:23liked or something like that.

00:36:24There are a few in there and in there a Karpathy too, who then also sits on the advisory board and

00:36:27looks at these topics again and also says, are these topics current, and looks for

00:36:31new connections, simply.

00:36:33So that means this vault in the background, this Second Brain in the background also has

00:36:37a kind of, without wanting to push the term too hard here, a sort of

00:36:41evolution of its own.

00:36:42So it grows along without me doing anything.

00:36:45So it recognises things too, they're built in, that I, when I don't react to certain messages

00:36:50I've passed it on to it via skills, so that bit by bit these topics then

00:36:57apparently aren't that important to me any more. I then try to set it up so

00:37:01that I say, but still have a look once a weekend, then there's another

00:37:05other skill that runs across it, that then has another look and says,

00:37:08wow, this topic is still so hot online and Jens still hasn't

00:37:13picked it up. Even though Jens apparently wants to slowly forget it or simply isn't attentive enough

00:37:19to my activation, I'll push it up one more time. If he then says no, then it's gone as well.

00:37:24So I've built in mechanisms like that. And that's obviously a massive advantage. I've

00:37:29also built in more things, like a guard that watches, because now and then I do want

00:37:32to access it with my OpenClaw as well and also with the online service in

00:37:38real-time working, so that I say, the guard also checks whether there's information in there

00:37:43that's okay for a public API or for my private API at Claude, which I then pay for,

00:37:50where I don't want to send everything up any more either. So basically it also

00:37:53checks what comes out of the Second Brain. And I think, in all the things

00:37:58I've just described, there are still a lot of topics in there that you can

00:38:02sort of derive, if we get to the topic, shall I build the bridge? Happy to hand to Mark or

00:38:08to Cornelius, one of the two of you, feel free to answer in that direction, how do we actually build something like that

00:38:12in a company context in future? I like that we've come back to the

00:38:16funding. I would have extended the bridge along the lines of,

00:38:21what advice do we give people to take away? But I'll save that bridge

00:38:25for the closing and I'll look towards Cornelius first after all,

00:38:30what do we do now? So we've all done a lot on the topic, yes, Second Brain and what

00:38:34have you. What do we do now in the company? What would be a good step

00:38:39that you could recommend in whichever company, do we have a bit of that? So I'll go one

00:38:44step very briefly back to Jens and then see that I make my way back into the company.

00:38:48We'll help you onto you. What I here for myself personally...

00:38:52To a company, right? To a company. What I've found for myself

00:38:59is actually that the systems often work reasonably well, but when I look at it in detail,

00:39:07quite a lot of mistakes do happen in places, and of course they propagate.

00:39:15So everything it auto-distils wrongly in the knowledge, or which derivation it makes, or also

00:39:22in part which things it comes up with, how it curates it, it makes a lot of mistakes

00:39:28And it makes a lot of mistakes, because I'm lazy too.

00:39:32It comes up with a good concept that sounds very good.

00:39:34And I think to myself, well, it has thought up something, how it

00:39:38scores and rates the whole thing, and it'll be fine.

00:39:42And then it builds that and then I also think, well, I do have a feature workflow

00:39:47and it does test driven design and it uses Codex for the review

00:39:50and it'll all be clean if it goes through.

00:39:53And then at some point you still do the check anyway and see

00:39:58whether it really does those things, or something strikes you, like just now in your graph model, that

00:40:02it somehow doesn't look good in some way, and then you let it think about it a bit longer,

00:40:07how the things actually work, and what comes out is, it made quite a

00:40:12lot of rubbish.

00:40:13And I think this, well at least at the moment I think, curation can't be fully

00:40:21handed off, and I think that will be the job as well, that we go in there

00:40:27and do a lot. And on the other hand I also think, well, if I save 90% of the

00:40:32time and it gets 30% wrong, I still have a big benefit. So

00:40:37I think this equation is right too. It doesn't have to be a perfect system. I get

00:40:41considerably more things that stimulate me in a good way and get me thinking

00:40:46than I would have had before, when I simply have to listen to things

00:40:50like that. But I do believe that curation is exactly the point that's

00:40:54super important for the corporate context. So that you think very carefully about which

00:40:59knowledge do I actually want to pass on to whom and make available in a skill, one that then

00:41:05has very high quality, is free of contradictions, maybe contains method knowledge, standard operating

00:41:11procedures, so business processes that you've defined, and so on

00:41:16and so forth. So things that fundamentally raise the quality level of the work. But

00:41:22the path there is one you have to shape very actively and with human influence.

00:41:29If I move on to my second question now,

00:41:33it's for the two of you then.

00:41:35You can work out between you who starts first.

00:41:38So what do you recommend now?

00:41:39Well, everyone who's stuck it out this far has to get a treat now.

00:41:44They have to get a treat, because they've understood by now,

00:41:49okay, maybe I don't want to explain everything from scratch to every model and every provider.

00:41:55I want to hold on to my interest, the one I have, somehow in the machine.

00:42:01I just want to start with the topic of Second Brain.

00:42:07So what are the three minutes in West, just to stake things out a bit,

00:42:13so that you can get started.

00:42:15a first step in that direction. Who wants to start?

00:42:21Yes, I'm still working out how that goes, that we both think about it although we're in different

00:42:25places. No, maybe that's a Second Brain thing too. Cornelius

00:42:28and I are already a virtual Second Brain, which would make it easier for us to work out.

00:42:32To work out together which of us starts. But I'll just do it now. I

00:42:35I think if you're starting now to think about building a Second Brain, then

00:42:45you really should, I think, if you haven't done it at all before, then

00:42:49I'd start with skills, the way you described it.

00:42:52Because I think the topic of skills is a good way in, to file a good skill somewhere

00:42:57away.

00:42:58There I have to, or to produce a good skill, it's not about the

00:43:00filing at all.

00:43:01It's about me wanting a good outcome afterwards with this

00:43:04For that I have to think about which task this skill is meant to do, to be able to handle

00:43:10and which additional information this skill still needs to have.

00:43:14So, and there I'm already right in that corner where I have to think about, okay,

00:43:20which information, structured in which way, do I have to

00:43:24make available to this skill so it can do its job.

00:43:27For me that's the first step into that kind of Second Brain thinking, because

00:43:30this Second Brain, the way I've built it. It's huge. Of course you somehow have to

00:43:34slice it up again for individual tasks later on. I don't think I'd advise anyone

00:43:40to sit down and build a monster like that first. That's actually more of a saving game

00:43:46that I'm playing at that moment. It starts with skills. First see that it builds sensible

00:43:51skills, because I think the combinations of skills with files that it also stores

00:43:56for these skills in project folders, something else, but on the computers is a nice

00:44:03first stage of a mini Second Brain that you build very specifically for one skill, because

00:44:10the Second Brain is basically nothing else.

00:44:13Before I hand over to Cornelius, quickly from the control room, what I actually

00:44:19made totally, I should say, simple at that point, you did say with skills

00:44:24and SkillSafe with the chances, I can save them and send them somewhere else. I did open up this SkillSafe

00:44:30and what I then tell the thing, here take the project out of the repo and look at that

00:44:37project and store in that project the knowledge from the following folders, and then I give it

00:44:43a folder, meaning I give it the transcripts from WWDC or from Google I/O and say here

00:44:47give it that, and then I can ask questions against it. I think you could

00:44:52actually do that pretty easily, because then the prompt is, no matter whether you take

00:44:55Gemini there or, or, or ChatGPT here or, or Anthropic, take this Git project.

00:45:04Here's my knowledge, mill it all together, give me back the skill.

00:45:08Thank you.

00:45:09That's a good way in, I think, and now I'm curious what

00:45:14Cornelius recommends.

00:45:16Something completely different, I think.

00:45:19And it's nice.

00:45:20Bananas.

00:45:21Yeah, that's totally good stuff.

00:45:24But again, more of a personal thing.

00:45:27I think for me it was more this thing of,

00:45:30how do I cut down my media consumption too,

00:45:35how do I manage to spend less time, so to speak,

00:45:39having to inform myself, or having the feeling

00:45:41of informing myself.

00:45:42And I think the thing that really helps

00:45:46or actually is, is shifting the focus away from the question,

00:45:50what else could I know? Or what else do I want to want to know? Towards, what am I actually

00:45:56working on right now? What do I actually want to achieve? And that already creates a focal point,

00:46:02to say, okay, which knowledge is even relevant at all. And then I've actually got

00:46:07a way in. I mean, for me it was mainly about this whole topic of curation really, and all

00:46:13sorts of content coming in. And I think it isn't that important to build yourself a

00:46:18perfect system that afterwards maybe just piles up a whole lot of rubbish,

00:46:25but rather that you get in with a different question, namely where do I want to go,

00:46:29what do I want to achieve, what do I want to build and what do I need for that. And then

00:46:33maybe not that much knowledge is needed at all, or it's a targeted search. And then

00:46:39you can build up your brain out of that bit by bit. That's what I'd start with.

00:46:44Okay, I think, good point from Nekonewos, I'd say, well, we've always said it,

00:46:53starting always helps, AI topic. I think just get on with it, so that again as a tip

00:46:59from my side. I don't think you can break much there. It's somehow also

00:47:03odd. We've got these super models that do everything with it, that can generate videos and

00:47:08read them out. And on the other hand, Second Brains at the moment are just

00:47:13sharing plain text. And that, I think, is on the other hand also nice again, because it's totally

00:47:16accessible, because everyone has at least once, back in their school days, tried

00:47:21to write a structured text and nothing read-in. Yeah, to read as well. It's nothing other than

00:47:28that, really, what we're doing here right now. We put, we try, to somehow store data in a structured way

00:47:33so that they're easy to understand in the first place, by the agents as well as for us, for the different use scenarios.

00:47:37That won't be the end of the line.

00:47:41So these things will keep evolving, in my opinion.

00:47:43We didn't really get to the topic of enterprise, enterprise brains today,

00:47:49a number of a thousand Second Brains that will interact with each other in future, or how that

00:47:54then works.

00:47:55We should maybe make up for that in another episode.

00:47:57So Cornelius, there's maybe the first invitation to you already, that you're welcome

00:48:01to be a guest again.

00:48:02We'll do a second episode on it and take a look at how basically

00:48:05these brains can interact with each other.

00:48:08But, as always, just try it out. There's nobody out there either who somehow says,

00:48:15this is plan A and who's doing it right, and accordingly, I think we managed

00:48:20to discuss and show a few things again today maybe, how you can do it,

00:48:26and also to show how you can't do it. If you think back to my

00:48:30story with the computer game, installing it, you don't necessarily have to use an LMV for that.

00:48:34I could have googled things there, probably. But never mind, I got

00:48:39caught up in the ride and then couldn't get out of it any more, because I also had no idea

00:48:42what the thing isn't doing on my machine at all. But never mind. And then you just give the AI

00:48:47full control over Jesus's bed. That can go wrong, right? That can go wrong. What's

00:48:53that Kassel office behind you doing, eh? Yeah, thank God nothing happened. Yeah,

00:49:00Yeah, but I think, Mark, do you want to, as so often, you always do that very nicely too, I have to say,

00:49:05we're somehow in a good mood and nice to each other, we have to somehow maybe again...

00:49:09I've already got a certain feeling of pressure in me, the way you're handing that to me.

00:49:16So from my side too, Cornelius, yes, even if one or two people maybe thought today,

00:49:21ah yeah, yeah, yeah, yeah, yeah, a lot of terms went out over the airwaves today that I first have to go and google.

00:49:26I'd say get started, try out the thing with the skills, I think that's

00:49:31actually very approachable, otherwise Notion, Obsidian, just, the main thing is

00:49:35to do it, and one thing also becomes clear to you, it feels like, once you start

00:49:42dealing with these primal ways of storing knowledge, Markdown if you like, you notice

00:49:49what you've actually lost in the days of, no bashing against

00:49:53Microsoft-year, but Word-Excel-Powerpoint, where you busy yourself immortalising information through bold type and

00:49:59the colour choice of a column, where you then think, yeah, honestly,

00:50:06there's another way to do it. Oddly enough, more and more people are busy right now with how

00:50:10you, in classic text files, with a hashtag. We used to call that a pig pen,

00:50:15but I don't think you can really say it like that these days any more, because they don't

00:50:19For headings and structuring elements and who knows what on the page, have a look at

00:50:23Markdown, look at how your information gets stored in there, look at how you might

00:50:28get transcripts from YouTube, podcasts or Professor Announcement or whatever interests you

00:50:34into there, and with that I'd say we're closing the first chapter of the Second

00:50:39Brain, coming to another episode surely soon again with the lovely Cornelius, if

00:50:45And we say goodbye in this episode, sending you off into the night with thoughts about expanding your own knowledge with the help of Mark Laundertein.

00:50:55With that in mind, see you soon, ciao.

00:50:58Ciao.

00:51:30SWR 2020