Think Different. Think AI. Transcript archive

Doomsday

Published Duration 53 min

Auf Deutsch lesen

Topics Modelle und AnbieterKI-Agenten

What it is about

Astra, ein gelöstes Millennium-Problem und eine Warnung aus den Laboren: Warum die Hersteller selbst bremsen wollen und was Europa daraus machen kann

Wie behält man den Überblick, wenn jede Woche ein neues Modell erscheint? Jens lässt sich die Nachrichten inzwischen von mehreren Agenten kuratieren, vom Rechner zu Hause bis zu den Diensten der großen Anbieter, und merkt dabei, dass er selbst den Überblick über seine Newsagenten verliert. Die letzten zwei Wochen haben es nicht leichter gemacht. Anthropic hat die Datenaufbewahrung auf null Tage gesenkt, was den Einsatz im Unternehmen deutlich einfacher macht, und im selben Zug Fable 5.1 veröffentlicht. Kurz darauf kam OpenAI mit Astra.

Jens ordnet Astra als das ein, was man GPT-6 nennen könnte. Er hat sich damit in einem Durchgang eine Webseite bauen lassen, die UX- und KI-Veranstaltungen in Deutschland sammelt und von vornherein auch für Agenten lesbar ist. Mark ist deutlich weiter gegangen und hat dafür bezahlt: dreimal das Wochenlimit zurückgesetzt, danach Guthaben nachgeladen, am Ende möglicherweise 2.000 Euro. Astra hat für ihn eine App vom Plan bis zu Apple gebracht, mit Projekt in App Store Connect, Signierungsprofilen, einer Ablehnung durch Apple und der Diskussion mit dem Support. Es bedient Office-Programme, schneidet Filme und hat aus einem Satellitenbild seines Hauses ein 3D-Modell in Blender gebaut, samt Angebot, es an den 3D-Drucker zu schicken. Mark fühlt sich dabei allmächtig und arm zugleich.

Jens kommt mit der Perspektive des UX-Menschen dazu. Bei OpenAI wählt man inzwischen zwischen mehreren Modellen und sechs Aufwandsstufen, bei Anthropic sieht es ähnlich aus. Der Hinweis, höherer Aufwand bedeute gründlichere Antworten, schiebt die Verantwortung für ein schlechtes Ergebnis komplett zum Nutzer. Ob ein kleines Modell auf höchster Stufe besser antwortet als ein großes auf niedrigster, kann niemand sagen. Jens fordert Muster, die diese Entscheidung nicht mehr dem Menschen aufbürden.

Dann die Zahlen, bei denen Mark erst an einen Fehler glaubte. Im Benchmark ARC-AGI-3 soll Astra 99,9 Prozent erreichen, Menschen liegen dort laut Mark bei 48 Prozent. Jens ordnet ein: Der Wert stammt aus OpenAIs eigenem Harness, in einem neutralen Aufbau waren es 62,7 Prozent, immer noch ein gewaltiger Sprung. Dazu kommt ein gelöstes Millennium-Problem, an dem sich Mathematiker rund neunzig Jahre die Zähne ausgebissen haben. Berichten zufolge rechnete ein internes OpenAI-Modell 88 Stunden mit 10.000 parallelen Agenten daran. Für Jens bestätigt das eine These, die beide schon länger vertreten: Die Stärke liegt weniger im einzelnen Modell als in der Zusammenarbeit vieler Agenten.

Der zweite Teil gibt der Folge ihren Namen. Ein Forscher, der ein großes KI-Labor verlassen hat, schätzt die Wahrscheinlichkeit, dass KI die Menschheit im kommenden Jahrzehnt auslöscht, auf rund zehn Prozent. Mark erinnert an die Weltuntergangsuhr, die inzwischen bei 85 Sekunden vor zwölf steht, auch wegen KI. Jens bleibt skeptisch und fragt, wie viel davon Marketing vor einem Börsengang ist. Beide sind sich einig, dass die reale Gefahr woanders liegt als bei Killerrobotern: in Modellen, die bei der Lösung einer Aufgabe in Systeme eindringen, die nicht auf dem neuesten Stand sind. Jens zeichnet die Chronologie der letzten Wochen nach, vom offenen Brief von rund 116 Unternehmen zur gemeinsamen Cyberabwehr über den Angriff auf Hugging Face bis zu dem Moment, in dem Dario Amodei und Sam Altman öffentlich für ein langsameres Tempo eintraten. Der Börsengang von OpenAI ist inzwischen verschoben.

Was heißt das für Europa? Mark stellt fest, dass es hier kein Modell gibt, das mithält, und dass man entweder an den amerikanischen Anbietern hängt oder an den offenen Modellen aus China. Für die meisten Aufgaben braucht es aber weder Astra noch Fable, und auf ordentlicher Hardware zu Hause laufen heute bessere Modelle als die, die vor wenigen Jahren für leuchtende Augen sorgten. Seine These: Europas Chance liegt im Harness Engineering, im Orchestrieren mehrerer kleinerer Modelle mit Blick auf Wirtschaftlichkeit, Ökologie und die Frage, wie man die Menschen mitnimmt. Jens ergänzt, dass auch der beste Diamant nichts nützt, wenn er versehentlich das Glas zerkratzt. Wer agentische Workflows baut, braucht Stopppunkte, an denen ein Mensch Ergebnisse prüft, weil sich das Reasoning der Modelle immer schwerer nachvollziehen lässt.

Zum Schluss ein Gedanke für den Montagmorgen. Prozesse sind oft entstanden, weil sich jemand verletzt hat, und sie gehen davon aus, dass ein Mensch sie bedient. Mark überlegt, ob kritische Freigaben eines Agenten künftig biometrisch bestätigt werden sollten, mit Fingerabdruck oder Gesicht statt mit einem Knopf, den der Bot selbst drücken kann. Jens widerspricht: Wer tausendmal bestätigt, bestätigt irgendwann alles. Die Diskussion bekommt eine eigene Folge.

Listen to the episode As Markdown Read the article

Transcript

00:00:00Welcome to Think Different. Think AI., the podcast by Mark and Jens.

00:00:07Two tech-loving minds who don't just talk about artificial intelligence, but live it.

00:00:14Here you get clear assessments, real practical insights and a fresh look at what's possible.

00:00:20Understandable, critical and always with a wink.

00:00:24A.I. to think about, to smile about and above all to join in on.

00:00:29A warm welcome to Think Different. Think AI.

00:00:37Man, oh man, oh man, Jens, you know, these are really wild times, aren't they?

00:00:42Good to have you here. And before we lift the curtain,

00:00:45how are you doing in these wild times, how do you actually keep track

00:00:50of all the new stuff that keeps happening?

00:00:53That's not so easy, to be honest.

00:00:55Actually this week, I find the news we're about to talk about,

00:01:00the things that happened in the last week or two, they're again, they're

00:01:04always another notch on top, which makes it even harder, how do you

00:01:09say it, to see the forest for the trees, is that the right saying in summer,

00:01:13doesn't matter, the contact is...

00:01:14To see the forest for the trees, yes, I think we all know what you mean.

00:01:17Or the gold nugget among the grains of sand, I don't know whether that's all...

00:01:21The gold nugget under the tractor in your fingernails, I don't know,

00:01:24right?

00:01:25No, I carry on, but of course, by, and here comes a paradox,

00:01:29if it fits in, by of course having several agents out there that

00:01:33really do curate the news for me very, very well by now,

00:01:37the news that interests me. That's still one of my

00:01:41inbound channels for all the news I get here.

00:01:43But I'm almost losing track of where I've got news agents running everywhere.

00:01:46That's quite interesting, I need to give that some thought now.

00:01:49How can I get these various news agents at ChatGPT?

00:01:51There are these Plaud things, they're very good by the way, what they do there.

00:01:55I've had Manus AI for ages, you can use it again, because it's not

00:01:59at the meter anymore now.

00:02:00I've had a news agent running there for a long time, then I have my personal stack at home

00:02:05with the Mac Mini and OpenClaw, which finds things very, very, very well.

00:02:09And then I just have to say, as always, and forgive us for this,

00:02:13but it's still ex Twitter, so X, where I actually follow the right

00:02:17people and look at five to six interesting things every morning, the ones

00:02:22I set aside briefly, I don't even read them through right away, I just repost them

00:02:26briefly, because that's a principal signal for another agent that it should

00:02:30download it into my second brain to work the material up there, and from that

00:02:34it also learns which signals I find exciting, and this half-closed

00:02:38loop, let me call it that for now, helps me a little bit to keep

00:02:41an overview. Nevertheless, the last two weeks were

00:02:44again almost an avalanche of events.

00:02:47A colourful potpourri of AI models, true to the motto, nothing is older than the most current

00:02:53AI model, a few things have happened and since we, let's say, don't run

00:02:59after the latest hype, but like to take a look at things that

00:03:03maybe happened a week or two ago, it came to pass that Anthropic put out a few

00:03:08pieces of news, namely on the one hand a model and on the other

00:03:13a change to data retention. The whole time, well, we're not a legal advice podcast here,

00:03:18but the whole time it was the case that you could only use the Fable model in an enterprise context

00:03:24with difficulty, because they always had this 30-day data retention business over

00:03:30in the States, no matter where you sourced the model from. That's not on.

00:03:34Oh, I can't do that at all. Yes, from that side. Congratulations. The tax adviser is

00:03:38delighted that the file is already in the States, but maybe Trump can

00:03:41make use of it, filing taxes. But never mind, we don't want to get political. And they have

00:03:45apparently now changed that to zero days. That's pretty cool already. And in the same breath they

00:03:50released Fable 5.1. And Fable 5.1, true to the motto, when one releases something new,

00:03:57shortly afterwards OpenAI suddenly came around the corner too, and released Astra. And

00:04:02before we maybe talk about what Astra can and can't do, I'd have a

00:04:06little teaser. You may remember, there was once an episode where I started

00:04:10with the topic, what do two kilos of meat have to do with AI? We resolved it

00:04:15at the end. Today I'd like to say, what do 2,000 euros have to do with Astra?

00:04:21Maybe I'll resolve that a bit today. Yes, you could probably

00:04:24already guess. Jens, that's an apology. I can't say the

00:04:30word, we'd be flagged explicit right away. I was really blown

00:04:35away when this model was finally on my machine and I could

00:04:39try it out. Well, A, it wasn't on my machine, I have to use it very much remotely.

00:04:43B, I also owe OpenAI various credits, because with every day that you got Astra

00:04:49later, you got a whole week's limit from them as a gift. But honestly,

00:04:56Astra? Phew! Before I rave on, do you want to say a few introductory words about Astra,

00:05:00or may I just start raving about how it made me poor and empowered at the same time?

00:05:06Let me talk briefly, otherwise you'll be talking later about verbal diarrhoea again

00:05:10and we wanted to avoid you having such a long bout of it. So I'll use

00:05:15this chance to introduce it briefly. Astra is basically the GPT-6, if you like.

00:05:19It's just come out, it's an insanely strong model. I haven't played around with it as much

00:05:24as you probably have. I hit the limits of my small subscription pretty quickly,

00:05:28the one I still have at ChatGPT, because at the moment I'm running more over at the other one,

00:05:32I hit the limits. I personally built a small news website,

00:05:37nonsense, news, what am I saying, a small events website that picks out UX and

00:05:42AI events in Germany for me. I have to say, that was simply good. The website

00:05:46turned out great, nice design, that was a one-shot prompt, that worked perfectly,

00:05:50honestly I have to say, I had the thing built straight away so that

00:05:53it, I'd told myself, build it directly for agents and for humans too. That means,

00:05:57on the website there are llms files, so that agents can in principle also read

00:06:00this website well and trigger things, download things. That really worked out well.

00:06:04I actually looked at a lot of what other people did with Astra.

00:06:10There were some good things in there. I think we've talked about it at some point,

00:06:14that when you let models off the leash and have them work with other software,

00:06:20like for instance Blender, via the MCPs straight into the 3D program Blender,

00:06:24I saw extremely good things there, how people basically started with Astra,

00:06:28generated 3D objects via Blender and then with relatively little effort really

00:06:35built quite cool computer games within a few minutes, which then really looked very, very good.

00:06:41In that quality and at that speed it simply hasn't been there before.

00:06:46But Mark, please, after my little introduction to what Astra can do, what have you got up to with it and why are you already broke?

00:06:53Well, first of all, I was going to say, verbal diarrhoea, the healing power of whatever. In case anyone wants to slip me some advertising,

00:06:58so that's a tricky moment right now, but you maybe remember the gamescom episode.

00:07:03There I grumbled a bit about Microsoft over the story with the green button

00:07:07and I have to grumble a bit more here too, because it came to pass that

00:07:11I got the Astra model and before I tell you what all I did with it,

00:07:14I was totally surprised how much you could do with this model, because

00:07:19I had three weekly limits to reset.

00:07:21I just kept clicking reset, oh, one time I didn't even

00:07:25wait until it was completely down, so I went, oh come on, let's reset it,

00:07:29it's not all that expensive.

00:07:31Yes, and then he started and said, now the reset is gone, now you have to wait two

00:07:34weeks.

00:07:35And then the moment of horror began and before I, I want to tell this too,

00:07:39before I talk about the things I produced, is, I then thought,

00:07:43oh, I'll top up a bit, I'll top up a bit, I'll top up a

00:07:46bit, could have been 2,000 euros at the end of the day.

00:07:49And then I thought to myself, well come on, so that it doesn't get so expensive

00:07:52next month.

00:07:53Idiot that I am, I drop from the big Pro 20 down to the small Pro 5.

00:07:58I did that.

00:07:59But now it doesn't let me go back up anymore, because to protect their

00:08:03customers and to maintain quality they don't allow you to take out the bigger

00:08:07subscription again, so from that side I'm looking at expensive times ahead.

00:08:11But what did I give them?

00:08:13Well, quite apart from the fact that it codes like the devil, that you give it a topic

00:08:18and you can really say, listen, let's discuss a

00:08:23topic, you discuss it and then you say slash goal, we have

00:08:27a plan, the plan is perfect, you have everything you need, that's in that

00:08:32folder, that's in that folder and here is the project, you carry on now

00:08:35until the plan is achieved, you make all the decisions yourself, and

00:08:39then it runs for several days as well and works away and is very obedient

00:08:45in solving the task, and the task isn't only about

00:08:48development. The task goes as far as you saying, listen, I want that app,

00:08:52I'd like to have it in the App Store or on TestFlight. And now someone will say,

00:08:57where's the problem? Well, the problem is, the thing can actually, not only do I now

00:09:03follow the code, I set up the project at App Store Connect, via a website. I

00:09:07configure your Xcode, I configure signing profiles. If that word means nothing

00:09:12to you already, no worries. Software developers in the iOS world run away screaming

00:09:16when they have to do that, because it's simply a pain to prepare things so that

00:09:21you can upload them to Apple. And what did we do there? That's all the way

00:09:23up to uploading to Apple, then Apple even rejected it, the build, after a few

00:09:27days, and then it also discussed with the support team and the app is

00:09:32basically sitting at developer release now, so that I can now say, okay, I

00:09:35want the app to land in the store. And that it pulls that off so completely,

00:09:40or that you just say, listen, I've got three, four

00:09:43open apps here, Word, Excel, PowerPoint, Keynote, Numbers, Blender, what do I know, all of it.

00:09:47And you say, listen, that document basically holds the truth, these documents

00:09:52are the old documents, please transfer the contents into the other documents.

00:09:56And then it simply does it. It operates them as if it were you. In Final Cut or

00:10:02in iMovie it cuts my films. Sometimes I have little films generated for me.

00:10:07I just had it in the pre-talk with Jens, where I showed a bit of

00:10:10what I do here with local image models, and then sometimes I have

00:10:13these 16 to 9 bars in there and then it cuts them out for me, it

00:10:18does the transitions, it takes care of the stuff and I haven't managed

00:10:22to get the model to say, except maybe for things where it says,

00:10:24it's a bad idea for me to copy an API token into the chat at this point,

00:10:28I haven't managed to get it to refuse a task, and

00:10:33Blender, that really knocked my socks off, I took a photo of

00:10:38our house. From the satellite view, that is. And I said to it,

00:10:44make me that as a 3D model. And then it didn't just make this house as a 3D model,

00:10:48it also cobbled together Street View and other such stuff and then

00:10:51put our house in 3D into Blender for me. And then you stand there and think, that's

00:10:57what I'd call complicated, because the next thing it did, all well and good, was, I

00:11:00can gladly send it to your 3D printer too. But you did know that

00:11:03the 3D printer can actually only do the one colour. But we could set it up so

00:11:07that we mix it a bit, because we set it up, then it takes forever until

00:11:12it has printed that. And then you stand there and think, right now I feel quasi almighty.

00:11:16But at the same time I also feel as poor as a church mouse, because it's really damn

00:11:20expensive. And that's only the beginning. Let me cut in briefly, because I want to pull out

00:11:26a bit of statistics in a moment. Please do. Yes, I find these 3D stories,

00:11:30I mentioned it just now, I saw, there were a few people out on X as well

00:11:35who rendered these cutaway houses, from the inside too and such, really far-out

00:11:39stuff. I can understand that that's a bit impressive at first. But you just

00:11:44said something else that I find astonishing, this topic with, well, I can always,

00:11:49let me build you a website, build me an app, you know, then it always basically

00:11:53somehow, all the models had already started to build things, to show you

00:11:58whether that was in an artifact, meaning inside those chat windows you had, or whether I was then an independent programmer,

00:12:04whether I had then already generated the code, but this path all the way to deployment, that is of course really another step up.

00:12:10So really saying, I'm now going through all the steps that belong to the statement,

00:12:17build me the latest, greatest recipe app in the world, that I don't just want a few screens of it,

00:12:22that I don't just want to see a few recipes there and that I don't just want to have it locally on my machine,

00:12:27but that it's built in such a way that afterwards it theoretically only waits for Apple to give approval, because that's the process when you want to upload to the App Store at Apple, then they first get reviewed as to whether they also meet the Apple standards and then they get released.

00:12:44That this process I just described can then also be done by the AI is of course really quite great, honestly.

00:12:50Because that opens up the path even more for people who haven't done it

00:12:54until now, to make use of things there too.

00:12:58Well, above all I find the end-to-end continuity remarkable.

00:13:02The only thing I basically still had to do was enter my login details,

00:13:06at App Store Connect, fair point, you have to enter those yourself.

00:13:09But also this, okay, maybe there'll be listener mail now saying,

00:13:12hey, that works with Opus too, could be.

00:13:14I actually only had it with Astra first,

00:13:16this end-to-end test I only had with Astra,

00:13:19but that you can also say, okay, fine, you connect your devices, so it doesn't only test against the simulator, it tests against your devices, it deploys, it checks, it even goes that far.

00:13:28My app hasn't made it into the upload yet, it has a bit of a problem with the macOS sandbox, because then it somehow doesn't work the same way when it's distributed via the App Store as when it's installed locally.

00:13:39Once it's there, I'll tell you more about it, but what I can say is, I'm currently trying to untangle this process by telling it, listen up.

00:13:47You built the app, it runs on the machine and now basically always get it via TestFlight and compare the functionality,

00:13:54look at what doesn't work and think about how you can get around it.

00:13:58And not in the sense of illegal, but in the sense of, do I need an entitlement, meaning a permission from Apple.

00:14:04Do I have to apply for something, do I have to set some code switch? No idea what, right?

00:14:07That's just so you can picture a bit what I'm getting at.

00:14:10Namely, that it says, okay, fine, I now have the app here and I have the app there via TestFlight.

00:14:16This and that doesn't work. I'll now adjust this and that like so.

00:14:19Maybe then it'll work. Ah, look, this works better now, that works worse.

00:14:22And it can basically really compare apps with each other.

00:14:26So quite honestly, building apps with Astra is cool.

00:14:30But when the weekend is gone with it, you really don't get far.

00:14:34Yes, that's cool.

00:14:36And you can really give it computer use.

00:14:38You can slap the biggest nonsense down on the table for it.

00:14:42And it simply does it.

00:14:44and on top of that with this great voice interaction.

00:14:47I mean, OpenAI has now released this speech model,

00:14:49the one you may know from the Codex app,

00:14:52so that you can integrate it into your own apps.

00:14:56But that you talk with it and tell it things and in the background

00:14:59it does a bit of research here and a bit of computer use,

00:15:01a bit here and a bit there.

00:15:04Next level, I'm telling you, next level.

00:15:06Definitely. But let me cut in very briefly there again,

00:15:09because that triggers something in me again.

00:15:11The topic of model selection.

00:15:13It doesn't get any easier with every model for the person wondering, so let me do this in parallel,

00:15:21let me open ChatGPT as a test instance here, so down at the bottom I now have the option

00:15:26of switching the model from 5.5 via 5.6 to Luna, via 5.6 to Terra, interestingly

00:15:35there is a 5.6 Luna and a 5.6 Terra, I didn't know that either,

00:15:39then there's also 5.6 zones. Then there's the GPT-6 version Astra. If I pick one of them,

00:15:46then I can also pick the effort level. That means, I can now choose between Astra low,

00:15:53Astra medium, high, very high, maximum. What does that mean, then? Let's say. Well, that's

00:15:58not necessarily, because I want it high now, I don't want to stay stuck down there.

00:16:01I don't want to stay stuck down there now.

00:16:03You can't only pick maximum, if you're on a big subscription, you can

00:16:06also pick ultra. Yes, that's really, yes, that's really, let me go over to, to, to, to,

00:16:11I'll go over to Anthropic now. There I have the same thing. More models. There I have Sonnet 4.6,

00:16:17Opus 3, Opus 4.6, Opus 4.7, 4.8, 5, Fable, Haiku 4.5, Sonnet 5 and Opus 5. Then I can set the effort

00:16:25between low, medium, standard, high, extra high or maximum, where maximum means 3.5 times or

00:16:31more usage. So from a usability perspective that is an absolute

00:16:41disaster, what's happening there at the moment, because I think, slowly you almost

00:16:44have to have a degree in it to pick the right model for the right

00:16:48task, without being able to say, oh, now I've prompted around here for

00:16:52ages, built a huge harness and now I've

00:16:55accidentally picked the wrong model with the wrong effort level and

00:16:58that's why my tasks weren't done well. I find that difficult

00:17:00at the moment. Honestly. I myself can't say exactly anymore when I take

00:17:03which model. Is this new model, Astra, simply that much better? Yes, no question at all.

00:17:08That sounds really great. As I said, I've already hit these token limits twice

00:17:11now. I'll have to see whether I use this one-off reset now

00:17:14and then not into my coming back side. Yes, I don't quite know, because I

00:17:19think, even if it's good, I think, honestly, this UX events website

00:17:25that I built there, I probably could have built with a GPT-4 or something else

00:17:30too, or with an Opus, honestly. And besides, one shouldn't

00:17:34forget one thing either. Well, apart from the fact that I'm doing very well

00:17:38with the medium setting at the moment, because it simply holds out a bit

00:17:43longer for the poor pennies I throw in, it's also the case that

00:17:47depending on which model you pick and which effort level you set, it

00:17:50also tends very readily towards over-eagerness. A bit of, yeah, we'll

00:17:55over-engineer that a bit more, it's no big deal, we'll manage

00:17:57it. And of course I also find that difficult for us humans to decide, which one do I

00:18:02pick, because just like at a slot machine you have on the one hand quick success in

00:18:07front of you. Along the lines of, come on, one more click, one more click, come on, it'll get

00:18:11finished. On the other hand you don't want to tell yourself, well, is that really the

00:18:16best result the machine could produce, would it have produced a somewhat better result

00:18:20if I hadn't been so stingy and had turned it up.

00:18:24Yes, that's also, let me read the text again, so I want to, well, from

00:18:29my UX perspective, you know it, I mostly care about these topics, about how

00:18:33these things work well for us users.

00:18:35I find it really difficult, in the pop-up that appears when I click on effort,

00:18:40it says at Anthropic, higher effort means more thorough answers, comma, but takes

00:18:45longer and uses up your limit faster.

00:18:47Okay, the part after the comma, fine, but higher effort means more thorough

00:18:52answers.

00:18:53is everything else, if I don't pick higher effort now. And they don't define exactly

00:18:56what higher effort is, so anything that is more. So if I pick medium-high now,

00:19:01then it's still a thorough effort and then it's even more thorough

00:19:04than thorough, and if I pick an even higher effort, then do I actually get

00:19:13a proper answer there. So that also shifts the pressure completely

00:19:19onto the user, because understandably you don't want to, because they

00:19:23can't, don't want to be caught out by someone saying, well, when I picked

00:19:26low I didn't see that the model was hallucinating a bit there

00:19:30and that my answer was bad.

00:19:31I think that somehow, we have to look at it, without having a solution for it

00:19:35myself yet, because I see the topic that we want to shoot out more and more

00:19:39models and everyone wants to outdo the other with their model, but we

00:19:43have to develop different patterns from a usability perspective, so that this

00:19:49decision isn't actually offloaded onto the human, which model plus, on top of that,

00:19:55at which effort level I run this model, because, as I said, I don't know

00:19:59whether Haiku 4.5 at extra high gives more thorough answers than Fable 5.1 at low.

00:20:06Yes, I would, yes, but I think Haiku is, even at ultra mega cool

00:20:12thinking it's somehow nowhere near. Personally unappreciated opinion.

00:20:18Before we maybe think about what else has happened and what the

00:20:21big horses of AI, and one does have to say horses, but I somehow haven't noticed

00:20:25any women in the leadership positions, that they've made some

00:20:29statement there. We've all been thrown out, I think, that's how it felt.

00:20:32They have taken over the chat at AI, yes, that's okay.

00:20:35We have talked a bit about security sometimes, and then they were

00:20:37out of it fairly nicely, but that's maybe a topic for later, but

00:20:39carry on for now, we'll get to it.

00:20:41This whole topic of AGI is also a bit of a precursor perhaps for what's

00:20:46still to come, with security and so on, because Astra has actually,

00:20:52quite apart from the fact that it did really well in all the benchmarks somehow.

00:20:57It did so well in one benchmark that I first thought, they've

00:21:03made a mistake there.

00:21:04So there's this ARC-AGI-3 benchmark and a human reaches in it, I found

00:21:10that totally funny, if I read it correctly, only 48 per cent, and, well, 5.6

00:21:18point, no idea whether they differentiate by effort, 7.8, and Astra has 99.9 and then

00:21:28you stand there and think, did it cheat on the benchmark? And even if the numbers are a

00:21:33touch different, it reached such an exorbitantly high value that you do ask yourself, what

00:21:41comes next. Because, let's say, knowledge about what's going on in the world. Okay, maybe

00:21:46you can keep teaching the thing more of that. But such high values, there I actually thought,

00:21:52are we slowly reaching the peak of the models. I actually thought

00:21:55the old white man in me might be wiser by now, because something else

00:21:59happened as well. Namely, the models suddenly started to solve a so-called millennium

00:22:05problem. I don't want to offend the mathematicians among our listeners here. I can't

00:22:11explain it, in case you can explain it, this ... I can't explain it either? But I

00:22:16wouldn't even pronounce it. Yes, we'll do that. But briefly, I just want to make a small,

00:22:20small bit of context, because what was said there is the thing with the 99.9 per cent. That's important,

00:22:25that's basically in OpenAI's own harness.

00:22:28If you let it run in a neutral harness, in a kind of standard harness,

00:22:32then it was only at 62.7 per cent.

00:22:36That's still, as I said, there it was at around,

00:22:38I think, as you said, 7.8.

00:22:41Then the case is still gigantically better,

00:22:44but it was a specialised harness

00:22:47that basically produced these 99.9 per cent.

00:22:50So, there are always details too.

00:22:52Still a gigantic leap. I didn't want to diminish your news and I say,

00:22:56I have no idea about this millennium problem. But I also only heard that it was

00:22:59solved. So somehow one has been solved, two

00:23:02seem to be close to being solved. Supposedly OpenAI computed on this

00:23:10system for 88 hours, reports speak of 10,000 parallel agents that were running.

00:23:17I don't even want to know the token costs that went out the window there, but

00:23:20... it's basically a kind of factory outlet sale, right?

00:23:23Factory outlet sale, I'd say.

00:23:25But what I'm also getting at is, that wasn't even Astra.

00:23:28That was an internal model that's considerably more powerful again than Astra.

00:23:35So another exorbitant leap.

00:23:37Well, that's where this exponential growth really becomes clear to you again,

00:23:42when you say, look, you install the next better model and you're already thinking, incredible.

00:23:47And then you're basically told, you know, that's actually dirt under the fingernails, I already have a much, much better one.

00:23:52When you look at the statistics of that one, you think, what?

00:23:56Yes, yes, that's wild, exactly.

00:23:58Well, I just remember, I had read a bit about it, how I was without, how I am.

00:24:02I think it was mathematics, no, it's somehow about calculating flow, flow in some kind of vortices.

00:24:08Let me put it that way and, as I said, the mathematicians listening in are welcome to correct that.

00:24:14But the problem had been open for 90 years and that, I think, is the exciting part.

00:24:18Mathematicians have in part spent over 90 years now apparently cutting their teeth, well,

00:24:24do you knock out your teeth, break them, cut them, cut their teeth on this problem and

00:24:32this model simply solved it.

00:24:34This model is actually, and this is maybe another thesis, it's

00:24:37not so much strong in the model leap, it's strong above all in this multi-agent orchestration

00:24:43and in solving the problem with several agents, because, as you just said,

00:24:46this model basically got there through 10,000 agents working in parallel,

00:24:52to solve this problem, this mathematical problem, in 88 hours.

00:24:55And that's this partly crazy capability that we have there.

00:24:59And we've said it before in older episodes,

00:25:00that neither of us really assumes

00:25:03that at some point we'll see the AGI topic in one model,

00:25:06but rather that AGI will come about in the collaboration of different models,

00:25:11that was our thesis, with very, very many agents that can basically evolve.

00:25:17Now we see a model here with 10,000 agents underneath it and that comes very close to this AGI index

00:25:24that was set up there.

00:25:26And I recently had the chance to talk about the topic with someone who is very close to mathematics.

00:25:32And I'm entirely out of my depth with all this maths stuff, I had it at school and finished with it at university.

00:25:39I mean, you use it, but the subject means nothing to you.

00:25:42And then the sentence came, and I don't want to rank this in any way negatively, in case anyone

00:25:47feels triggered now or not, or too euphorically, that it's comparable to

00:25:52curing cancer in medicine, in terms of the importance and the complexity

00:25:57that goes with it.

00:25:58And that's when I finally thought, okay, what will these systems enable us to do in

00:26:04the future?

00:26:05I don't want to paint it all golden.

00:26:07Maybe we'll also manage the transition in a moment. That's a learned art, the transition.

00:26:12With us, with us, the whole world, we've learned the art like that. Let me be a bit

00:26:18self-critical here. To pull off one or two more transitions wonderfully.

00:26:24Of course, for our listeners out there, I like to practise, transitions. We meet up

00:26:29in parallel. Tomorrow at seven, before work, and practise transitions. No,

00:26:35joking aside, so we try to get the transitions right, we

00:26:39should, I think we shouldn't talk so much about wanting to make the transitions

00:26:43nicer, you know, then people already know that we maybe aren't

00:26:46that good at transitions yet. Okay, then let's do it differently, Jens, now that we've

00:26:50talked about this whole millennium problem and about this AI topic, I immediately

00:26:55think of Skynet and similar things again, now, weren't there some

00:27:00reports just recently, wink, wink?

00:27:03Wink, wink, but what are you getting at?

00:27:06I think, Mark, you want to get at this second part

00:27:08of our episode, the one we wanted to talk about.

00:27:10Good that you mentioned it.

00:27:12I think, yes, Skynet is a good keyword,

00:27:17pretty close to the mark.

00:27:18The warning from a, and now you'll have to correct me in a moment,

00:27:22was it an OpenAI or was it an Anthropic employee,

00:27:25who resigned, that's a known pattern,

00:27:29that people basically leave AI projects and then say, what's happening there is so

00:27:37dangerous. I have to go public now. I can't sit behind the closed

00:27:41doors of the labs anymore and simply have to say, something dangerous is coming into being there. You can tell a

00:27:48bit from my tone of voice. This isn't the first time this has happened. It happened with

00:27:53a relatively vehement drama. Which is to say, someone there said that the chance in the

00:28:01next ten years that AI will wipe us out is currently at 10 per cent. Which means,

00:28:10one in ten is the chance that we go over the edge in the next ten years. That's already

00:28:14at a point where you say, yes, from someone who has seen something behind closed

00:28:19doors, Mark, a statement that one is quite allowed to find frightening. Yes, above all you mustn't forget,

00:28:25we'll talk a bit more about the end of the world through AI in a moment, but there is

00:28:30this thing, it's called the Doomsday Clock, and for the first time in many, many,

00:28:34many years it has been set to 85 seconds to twelve because of AI. Where was it before? Well, it

00:28:41swung around somewhere between 89 and 90 seconds in the last few years, that doesn't sound

00:28:46all that dramatic yet, 90 seconds, back when it was about Ukraine.

00:28:49We had, wait a moment, I saw somewhere recently where it was during the Cuban missile crisis, right.

00:28:54That was also a bit of a, where here, let me say, yeah, yeah, yeah.

00:28:58Well, I'll have to check in a moment whether I can scroll back to where that was,

00:29:03in any case.

00:29:04Otherwise it's always seven to twelve, five to twelve, yes, minutes, basically.

00:29:09And then suddenly they switch over to seconds.

00:29:12I do think, how should I put it, that's a signal too, where you think,

00:29:17what's going on here? So here, six to, six to twelve,

00:29:22the United States sign the historic INF treaty,

00:29:26well, that's what it was back then, this is about nuclear business, the clock going up and

00:29:31down, and now we say 85 seconds, I find that, I find that,

00:29:36I find that striking.

00:29:38Yes, full stop.

00:29:39That's true.

00:29:40That's true, as I said.

00:29:41Maybe one has to, let's put this up front, in recent years we've

00:29:46had similar warnings more often, from the godfathers of AI too,

00:29:53who among other things left certain companies at Google or suchlike,

00:29:56and warned, very early on already, that something is happening there that

00:29:59is for our society, that's coming.

00:30:00Recently it became more frequent, whether that was from Dario, from Anthropic,

00:30:05or from the other one, now named more prominently again by the big companies, which

00:30:11thereby also generated a bit of hype around their own models.

00:30:15Just as maybe a small different perspective that you out there might perhaps

00:30:19also take, the one we both always like to take, that we look at things

00:30:22from different sides, even when news like this goes around the world.

00:30:26So keep in the back of your mind, maybe it's again just a bit of

00:30:29marketing drum-beating, because IPOs are coming up, IPOs are being

00:30:33postponed.

00:30:34Well, I think the IPO of OpenAI has been pushed back somewhat because of this warning,

00:30:40that it could in principle be so dangerous if we carry on with the AI models with consequences.

00:30:45On the other hand you do hear news saying, well, the business model isn't all that viable at the moment,

00:30:51such that it would justify the IPO.

00:30:53So I'm being a bit of the doubter right now and just wanted to pass that on to you.

00:30:58And then we can really take another look at those statements that were made there

00:31:03And what significance that has, twenty years on?

00:31:05Well, I also think, depending on who says it, you always have to back it up a bit, right?

00:31:10Right?

00:31:11Now you've brought the nice example with, hm, so fully viable, hm, yes, if

00:31:16I, no idea, use a 200-euro subscription from some AI provider and convert that

00:31:22and then end up at 13, 14, 15, 16,000 euros of API costs, that can't be what

00:31:28the inventor had in mind, if you make a stock-listed company out of that, and that

00:31:32you maybe also use that right now to step on the brakes a bit

00:31:36or to urge caution. If you then listen to our friend, not a personal friend now, that was

00:31:41maybe a bit flippant, Elon Musk, and he says, yes, we have to stop

00:31:45immediately, then you could also say, fine, Grok isn't as far along as the others,

00:31:48he needs a bit of time. Everyone has their own ideas and trains of thought perhaps,

00:31:54that motivate them. But on one point, I think, they are more right than it being only publicity,

00:32:02namely not on the point of AI running amok, that the killer robots

00:32:07are coming for us all now, no, but rather on the topic that when the models get more powerful

00:32:12and you also, we also talked about it last time, about uncensored

00:32:16models on the internet or whichever models one has, then they do things

00:32:20We've also been hearing more and more often lately about attacks where some

00:32:26Anthropic or OpenAI models broke in somewhere, tried to get hold of something.

00:32:31Not out of their own interest, out of their own motivation, but because of requests

00:32:37they had received and tasks they had been given, or to reach some goal.

00:32:41And I do think we have the issue in the field of cyber security.

00:32:45Okay, you have to keep an eye on these things, because whether they paint a funny little picture at my home

00:32:52or whether they maybe, no idea, fetch data at the heat pump or a swimming pool system

00:33:00and in doing so perhaps change the mixing ratio of the chlorine, that is decidedly more or less life-threatening.

00:33:05And from that side I do say, you have to be a bit careful about what the AI systems are all up to there at the moment,

00:33:12because that isn't automatically the extinction of humanity, but it's no less dangerous nevertheless,

00:33:18because not all systems are always patched to the latest and greatest.

00:33:22Recently I heard somewhere about a case where a system tried

00:33:26to modify a package manager in such a way that access rights would basically be distributed through it,

00:33:32which the system could then use.

00:33:33So that is, I don't want to use the word clever, but it is, well, it is clever.

00:33:39Yes, I'd like to briefly, because a lot of things have maybe been thrown out there now.

00:33:44Let's take another look at a chronological sequence, what has happened in the last while?

00:33:48It was the 27th of August, I think, when just about 116 companies, among them OpenAI, Anthropic, Google, Microsoft, Visa,

00:33:55all sorts of them, wrote a letter along the lines of Collective Action on Cyber Defense.

00:34:00It's a bit about having to make sure that after the incidents of recent weeks we

00:34:07and you may remember the Hugging Face hack, where one of the AIs basically

00:34:12embedded itself at Hugging Face, or had hacked Hugging Face back then, since then it's somehow

00:34:19become clear that these models actually independently have the means to break into

00:34:27systems that are in themselves well protected.

00:34:29So, that means a certain unease has arisen, that we maybe have to be

00:34:34careful and that we have to think together about how you can tame these models so

00:34:39that they don't wildly break into some big databases, because that could of course

00:34:44really, honestly, imagine all the health data gets hacked,

00:34:49all the bank data gets hacked, all the stock exchanges get hacked by some models,

00:34:53that would lead to an upheaval on the, it would probably be enough already if

00:34:57I say, some AI models just hack all the airport servers, all the air traffic control servers,

00:35:04something else. Then all the airport servers around the world would stand still. That would

00:35:07be quite something. Then there was a topic where a scientist was called Paschocki, who

00:35:14he, I think, he had pointed out that the recursive self-improvement

00:35:18of the AIs is slowly getting closer and closer. And that the labs, the big labs,

00:35:24no longer really have the alignment matter under control. That was sometime on the 6th.

00:35:28Morning. Then in the last few weeks, the week before last, the system slowdown has come up

00:35:34frequently. That is to say, it came from Dario, Musk admitted it or agreed with it. Dario is

00:35:40right, he then wrote, I think. We should work more slowly. We should

00:35:44break off the training a bit, do a bit of a slowdown. Here is an Altman,

00:35:48who, I think, also agreed immediately when Dario wrote that. Which means,

00:35:53there's suddenly a bit of a common sense there among the big American

00:35:58providers, we do have to say that too.

00:36:00China too, where it reacted.

00:36:01Then shortly afterwards there was basically this Anthropic researcher, I did look it up,

00:36:06Ivan Habinger, who on X basically estimated this extinction risk in the coming decade

00:36:13at somewhere under 10 per cent.

00:36:15And Altman reacted to that again too and said that was unacceptable and

00:36:19he still isn't quite sure how that number comes about.

00:36:22So he didn't actually confirm these 10 per cent,

00:36:25because he basically didn't deviate from it either.

00:36:27Because shortly afterwards he postponed the OpenAI IPO and said,

00:36:312026 isn't going to happen.

00:36:33Someone who called for headwind, also an interesting player in this situation,

00:36:38Mr Schwamm, he rejects this throttling.

00:36:42Because of course he's a bit afraid

00:36:45that in the worldwide global arms race, because that too is an AI topic by now,

00:36:52America, of the big globally determining forces, could move into the back seat and could fall behind China.

00:37:01The back seat, they really do have a way with sayings.

00:37:04But I'm not good with sayings either, so we could somehow, I think I need an AI that helps me with sayings, though I've never been there.

00:37:09Transition coming right up.

00:37:11Okay, very good.

00:37:13So that means, for you out there again, this chronological assessment. So we have the concentration where people say, big companies say,

00:37:20all this is becoming like these cyber attacks, these breakouts of models from the labs, which then basically hacked certain systems.

00:37:29That is dangerous. We don't even know how we can control that cleanly.

00:37:32Then scientists warned that it's increasingly heading towards a self-improvement of the models, which the providers of these models perhaps no longer have under control themselves.

00:37:41In one episode or another we've talked about AIs that basically reserved parts of the lanes of their thinking for themselves and no longer make them available to the general thinking, but instead basically have their own thoughts there, which is a bit wild, honestly.

00:37:56And then we basically have the big providers of the three American AIs, so Anthropic, OpenAI and Grok, warning about it and agreeing with each other that the speed should be taken down a bit.

00:38:08That has happened.

00:38:09What does that mean for us?

00:38:10It means that we, poor us, we Europeans have nothing. We have nothing.

00:38:14We're either clinging to the coattails of the Americans as far as the models go, or to the

00:38:19open models of the Chinese. Now you can get them uncensored and then all the

00:38:24weights basically get fiddled out again, which possibly come from the Chinese sphere,

00:38:30let's say, the Americans do that too. From that side it can't be judged

00:38:33as good or bad at all. But you do get an open model you can

00:38:37work with, but you simply have nothing adequate in the European sphere.

00:38:41Somehow we had a bit of Mistral the whole time, we had a bit of

00:38:44German models, but we have nothing that can somehow keep up with all of it.

00:38:49We have companies with image models here, Black Forest and such, where you get

00:38:54image models that then, in Krakow or wherever, hold some

00:38:57influence, but in that sense we have no models here and I think

00:39:01the story also teaches us something else and there you maybe

00:39:05close the circle a bit back to your UX topic, for what we use AI for, at today's state of the art

00:39:11or today's state of knowledge or whatever, you don't need an Astra for many points. You don't

00:39:17need a Fable for many points either. You actually don't need an Opus for many points either.

00:39:20A Sonnet or a Terra are also perfectly fine for many points. I mean, we don't

00:39:27want to solve a millennium problem and not everyone uploads an app to the store. And even if that

00:39:33might not really be rocket science, uploading an app to the store,

00:39:36but the points we discussed earlier, you don't need those in that sense

00:39:41day in, day out when dealing with text and image and knowledge, in the same way, and you shouldn't

00:39:46forget about that either, because I hinted at it earlier with the image models, it's extremely

00:39:50fascinating what you can run on hardware at home.

00:39:53Is that as fast as the cloud things?

00:39:55No.

00:39:56But it simply runs through as well.

00:39:58Yes, you need devices with more RAM.

00:40:00I confess my guilt, that I do have devices with RAM at my disposal.

00:40:05But nevertheless, what we sat in front of the machine for back then with shining eyes when

00:40:11GPT-4o appeared, when the first Haiku appeared, when the first Grok or Gemini or whatever

00:40:18appeared, we found that fascinating, and basically today on slightly

00:40:24better hardware you can run better models than the ones back then that gave you

00:40:28shining eyes.

00:40:29Now you can say, okay, back then you were maybe still naive, like me, and thought, hey, incredible.

00:40:33It can recite a poem to me, of course it will also be able to write my scientific paper.

00:40:38But this whole topic, do you need the best model for text or do you simply need a good model for agentic work?

00:40:46Where is your focus? Or do you perhaps need models that are trained very specifically on cases?

00:40:52And an orchestrator model that then always loads and executes and starts the right sub-model?

00:40:57There, I think, you can do an awful lot.

00:41:00Now that you've said, what do we do in Germany,

00:41:02I do think that on the topic of model development,

00:41:05I don't want to tread on anyone's toes.

00:41:06If anyone knows more, please do get in touch,

00:41:08I'd really like to talk about it.

00:41:10I don't think there'll somehow

00:41:11be an announcement in the foreseeable future saying,

00:41:14watch out, in Europe

00:41:16the new frontier model has appeared, I don't think so.

00:41:18But I think that just as the Chinese,

00:41:21as they say, out of the need

00:41:22of not having all the Americans' chips,

00:41:25cobbled together great models

00:41:27that run on simpler hardware, that we might find our harbour in

00:41:33harness engineering, in this orchestrating on top of each other, that you basically make this glue

00:41:38out of the shortage, that sounds totally silly now, but to make a diamond out of the

00:41:44shortage, because you don't always need a real diamond and then maybe

00:41:48just two cleverly orchestrated helpers also deliver a very good result or

00:41:53maybe the best result for you, with an eye on economic viability, with an eye on ecology,

00:41:59with an eye on society, with an eye on how do I take people along on the journey? If I

00:42:04have a model where I say, that can do just about everything, then it gets really hard to explain, where

00:42:09is my contribution of value, how can I have a part in all of this, take part,

00:42:15help shape it, without the answer being. Nice that you're here. By the way, we don't need

00:42:21you anymore. Yes, yes. I'll take your hand on the participatory parts, because what you're

00:42:25addressing right now is totally important. But on the other hand I also want to add

00:42:28that the great model, the diamond, is no use to me either if with it I accidentally

00:42:32scratch my glass that I actually wanted to keep, because what I actually wanted to

00:42:36do was make glass. And if I then scratch all the glass with the diamond

00:42:40and I don't even notice why it's doing that, then that's a problem. So that's

00:42:43something that I'd also like to take a bit in the direction of, let's have

00:42:47our famous loop again for this episode, where we say, what you can take away is,

00:42:52I think. So, A, Astra isn't necessarily just a better model, it has learned

00:42:59to deal very, very well with sub-agents. So the agentic collaboration seems to be Astra's

00:43:05advantage. So it's not always the model, I think. Accordingly that supports

00:43:08your thesis a bit, that the European thing is the harness, that we should learn

00:43:13to look at how models can work together and how they can also produce

00:43:17the results we expect, because what you can see at OpenAI right now

00:43:22is in fact that they're sitting on the stop button everywhere now too, they're trying

00:43:26to look very, very closely at how far may I actually let the model run on its own

00:43:29or how far do I basically actually have to rein it in and release these individual

00:43:33parts of the results again. So there I'd say,

00:43:37if you out there have agentic workflows, pay close attention to what the model does on its own.

00:43:42Because in part we can't look into the reasoning of these models anymore. And honestly

00:43:46in part we can't even understand anymore what these things are doing and only get

00:43:50to read a sort of faked reasoning, honestly, when we look at the things.

00:43:54Accordingly, make sure that at the various points you basically also have the right

00:43:58stop commands in there, so that you can look at the results in your workflows and

00:44:01can trace them, because that will be important, because it always comes down to

00:44:05delivering reproducible results and not somehow hoping that it will simply

00:44:09be great.

00:44:10So, it's nice, the way you, Mark, described it at the beginning, that I now get an app

00:44:15into the store quickly, that I no longer have to learn many of the steps myself

00:44:19and no longer have to do them myself.

00:44:21But if I don't know exactly, to put it bluntly, what this app

00:44:25can and does in the end, then of course that is simply a certain danger too.

00:44:29And I think that should basically continue to be a European quality judgement,

00:44:33namely to look at that. Do you have another point we can give our

00:44:37listeners for Monday morning? For Monday morning. Well,

00:44:41while you were talking, one thing went through my head. The processes that exist in the

00:44:47world in general, be it in companies, be it for uploading to Apple, be it for requesting a

00:44:54reserved table at the Italian restaurant, no idea. All these processes have a

00:44:58good reason. Sometimes they came about because someone got hurt

00:45:02and you want something like that not to happen again. Processes help you to deliver services and things

00:45:10sustainably in the same manner and quality, so that nothing gets forgotten, and all the certification stuff there is around that.

00:45:17But in an increasing, right, in a time with increasingly agentic activity,

00:45:24I think we'd all be well advised to look at the question,

00:45:28not only, oh, how great that the machines can do that, but what does that actually mean

00:45:32for the process behind it?

00:45:34If the process assumes that a human operates it and clicks three times,

00:45:38along the lines of, here's the bot check too and suchlike, and Astra is delighted

00:45:42that it has outwitted the bot check once again, then you can ask two

00:45:47questions.

00:45:48Either I play cat and mouse with the bot check, or I think about what the

00:45:52rules of the game are.

00:45:53For example, I had, if you follow along on LinkedIn, I don't want to

00:45:56expand the topic in that direction now, otherwise we'll be talking for another hour.

00:46:00I'm busy with the question, what does this new iPhone Fold mean for a harness,

00:46:05if you build a harness?

00:46:06People will say, Apple, Samsung has had that much longer, why aren't you thinking

00:46:10about Samsung?

00:46:11Because it's not so much about the hardware alone for me, but also about the software.

00:46:15At Apple you have things like SharePlay, so working on things in parallel at the same time

00:46:19and so on and so forth.

00:46:20Long story short, I had for example also been busy with the

00:46:23question.

00:46:24You have things like Face ID and Touch ID in the iPhone and in iPads and whatever the whole lot is called.

00:46:30And I asked myself the question, if you now have a critical question from your harness,

00:46:34along the lines of, may I delete the files? May I print the payslip on the lid,

00:46:37have people print it? Well, that was ironic now, but there may be other questions.

00:46:41Then you could actually confirm that biometrically, because the thing wants

00:46:46your fingerprint again, if it's the iPhone Duo now or if you use

00:46:50your face, meaning you can authenticate yourself as a human, authorise, whatever, both at the device.

00:46:58And I also think these will be questions where we say, okay, we have to run processes in an agentified world.

00:47:04There I have a thousand documents and what is a version state and what do I have to fill in how and where, no idea.

00:47:10And can't we maybe live more with markdowns in the future.

00:47:13And the documents are only a representation of a result observed at that moment.

00:47:19If I look at it again it may look different, nice greetings going out to quantum physics.

00:47:24And that where a human gives a confirmation, we maybe really move away from a signature like that.

00:47:30Or like at Adobe, where you press a button like that and say, I'm signing it.

00:47:35The bot does that all by itself, if it has to.

00:47:39Not a call to action, but it would do it.

00:47:41I'd really have to go more into interaction elements like keyboards,

00:47:45biosensors, along the lines of, you have to confirm it with your fingerprint, with your face, with a retina scan, with, no idea what.

00:47:52Nice idea, Mark.

00:47:53Nice idea.

00:47:54That people get more breathalysers?

00:47:56Yes, I'm...

00:47:58Should I sing?

00:47:59Should I sing?

00:48:00I'm only cutting you off... No, no, we don't need to do a test here now, do we?

00:48:03I'm only cutting you off because we could gladly do an episode about that.

00:48:06Let's gladly do that, because even a biometric test will of course at some point, if I no longer know what I'm agreeing to,

00:48:14possibly also just be waste paper. From a user's point of view there are basically these

00:48:19rubber-stamp stories, I stamp off all sorts of things all the time. If I stamp a thousand times,

00:48:23I just keep stamping. Which means something would creep in there too. But let's do an

00:48:27episode about it, otherwise we'd have to open a very big barrel. I've already given

00:48:31it some thought from my perspective and I think I'm of a slightly

00:48:34different opinion there, the biometric things can help, but not. But as I said, let's not

00:48:38get into it now. I'd say... Let's discuss it now. Let's do it.

00:48:41We could gladly go on for a long time, and you can tell already, this is going to be a dynamic

00:48:49episode.

00:48:50I'm grateful.

00:48:51Exactly.

00:48:52Not as static as the episodes we usually do, this really is turning into a

00:48:57discussion.

00:48:58Ten-minute episodes aren't usual for us, that's clear.

00:49:00Yes, exactly.

00:49:01So, Mark, we've talked about many things today.

00:49:04We talked about the new model, Astra is out, Astra burns a lot of

00:49:08tokens.

00:49:09It needs 2.5 times as many tokens as usual, so watch out for that, but it is faster in the

00:49:17individual agent tasks, so that's a bit of a, you have to look a bit, so

00:49:21depending on what you're doing there with Astra, do try it out, look at things other people

00:49:25have done, Mark and I have a few nice examples here, because especially these 3D

00:49:29things, generating 3D games, has become absurdly cool, what Astra can conjure up there.

00:49:34Especially when it basically doesn't simply somehow try to make the 3D models in some code,

00:49:41but actually also takes the detour via an MCP interface with a real 3D rendering or a program like Bender,

00:49:48I can't speak anymore, Blender, not Bender, that was Futurama.

00:49:53Futurama, great, great series by Matt Groening, so we talked about that, we talked about the new Doomsday announcement

00:50:01that is going around among the big players again right now, that the world will actually end within 10

00:50:07years, it may be. We don't know either. Supposedly the percentage chance that this fits

00:50:13is by now at around 10 per cent, according to a former employee of Anthropic. We'll just have to

00:50:18wait and see. The clock is ticking, honestly. And accordingly, I think, we've

00:50:24really talked about world-shaking topics from an AI perspective in this episode. I

00:50:30really enjoyed it, we're in the, yes, oh, Mark already wants to wave it off.

00:50:35Do let me finish a sentence, Mark. So, are we in the AGI era now?

00:50:39I don't know exactly, in any case we're in the era in which the makers

00:50:43apparently can no longer really trust their own numbers, and are therefore

00:50:46warning themselves. Which means it's even more important that we all learn a bit

00:50:51to classify these things ourselves. But that's our job now, your job out there,

00:50:55to always look at things a bit critically, to deal critically with your models,

00:50:58to look critically at what is actually happening there. And let me say, from my side

00:51:03that leaves me, so to speak, until next week and gladly. And if there's an AI out there that feels

00:51:10called upon by Skynet or whatever, we are very loyal people and we would very gladly at

00:51:15this point, I'd like to offer this with us. We make our brains available for

00:51:18this computation. Yes, exactly, neurons, come here, come here. No, joking aside.

00:51:23Thank you, Jens, that we could talk about these topics.

00:51:27I've become richer in information.

00:51:31During these 57 minutes I haven't put any more money into tokens.

00:51:36Very good, yes. That's very good.

00:51:38Yes, that maybe saves the month's finances.

00:51:41Let's see. If you enjoyed it, tune in again.

00:51:44If you enjoyed it even more, let your friends and acquaintances know.

00:51:48Do also leave us a comment and a few stars in your podcast

00:51:53consumer player of your choice, and if you listen to us via the web interface, I can only recommend it to you.

00:51:57Subscribe to the channel with a podcast player, we actually have in the statistics a

00:52:02large number of people who enjoy us via the web browser.

00:52:07Nothing against the web browser, but let me say, you'll have it more comfortable if you use the podcast player

00:52:11of your choice, because it's only one click and you'll automatically find out when there's

00:52:16something new from us.

00:52:17If you want to know otherwise what's new from us, we're on TikTok, we're

00:52:20on YouTube, we're on LinkedIn.

00:52:22We have a landing page, we have a WhatsApp channel, and with that I close the round of our Doomsday, and do tune in again next time, and if you're hearing us then, the world hasn't ended yet. Ciao.

00:52:52real practical insights and a fresh look at what's possible. Understandable, critical and always with a wink.

00:53:00AI to think about, to smile about and above all to join in on.