Physical AI Will Be This Year's Christmas Gift

Share
Physical AI Will Be This Year's Christmas Gift
๐Ÿ“… September 11, 2026
๐Ÿค Bloom ร— Nebius ร— NVIDIA
๐ŸŽค Guan-ru Huang, Cloud Solutions Architect at Nebius ยท Dhruv Diddi, Physical AI at Nebius ยท Andy Lee, APAC Head of Physical AI and Inception at NVIDIA ยท Yu Been Park, CBO of Diden Robotics
๐Ÿ“ Woomul, Hapjeong, Seoul
๐ŸŽŸ๏ธ Event page

A round of applause for Nebius and NVIDIA, who built a great night in Seoul with Bloom.

A global AI hackathon has just begun in Korea. Around 150 AI builders from many industries gathered with Nebius around the theme of physical AI. Korea has been a manufacturing powerhouse for a long time, and the opportunity the new AI wave brings is right next door.

I do not think I will forget the moment I heard that this Christmas, children might be unwrapping robots under the tree. The change is happening in a blink. Let us keep building with Nebius and keep pushing the edge of what is possible.

Deep thanks to Dhruv Diddi, Guan-ru Huang, Alice Lin, and Daniel Colaianni of Nebius for backing us so solidly, and to Andy Lee of NVIDIA and Yu Been Park of Diden Robotics for the talks that gave us so much to think about. Thanks as well to our wonderful Bloom ambassadors Changwon Jeon, Minjae Lee, Jueon Kim, and Jungmin Kang, and to the whole Bloom community.

We wanted to at least beat Tokyo

The venue was Woomul, a cultural complex in Hapjeong. It runs about 660 square meters with an outdoor yard, indoor second and third floors, and a rooftop, and we rented the whole thing. Bloom usually gathers in lecture-hall type spaces, and the night before we had run an indoor event with another company, but this one opened a hackathon, so we picked something different.

Even against a Friday evening commute, people arrived early from the 5 p.m. check-in. We ate dinner first and said hello to whoever was next to us, then started at 5:30. When I asked who was at Bloom for the first time, about a quarter of the room raised a hand, and I also saw faces from the previous night.

Nebius is a Nasdaq-listed AI infrastructure company, and this hackathon is a global one it runs with NVIDIA. It started in Tokyo two days earlier and travels through Da Nang, Seoul, Kuala Lumpur, Singapore, and Taipei, then London, Paris, Berlin, New York, Toronto, and San Francisco, twenty cities in all. In the opening I said out loud exactly what I had been feeling while preparing.

Preparing this, I felt a little competitive. We should at least do better than Tokyo. We should at least do better than London.

Of the twenty cities, I wanted ours to be remembered as the coolest, the most attended, and the most widely shared. We opened a dedicated Discord room just for this hackathon.

Today is not a submission day, it is a starting line

This was not a night for finished work. The hackathon runs online, and the submission deadline is 10 a.m. Pacific on October 30, which is 2 a.m. Korean time on October 31. Anyone can enter even without attending the Seoul event, so we framed the night as a starting point for forming teams and getting the tools into your hands. The Luma page said outright that there would be no demos and no pitching.

To enter, your build has to run on Nebius Token Factory or Nebius AI Cloud, and use at least one NVIDIA open-source model. On the language side that means the Nemotron 3 family; on the physical AI side, Cosmos, GR00T, and Sonic. There are four tracks: coding and agentic engineering, for coding agents and developer tools that write, run, and test code in the Token Factory sandbox; apps and agents, for something a real person will actually use; personal AI, for an always-on assistant that keeps your data in your hands; and physical AI, for agents that sense and move in the real world through robotics, IoT, and on-device work.

Prizes total more than 50,000 dollars. There is 20,000 for the grand prize, 10,000 for second, and 6,000 for third, and the four track winners receive an NVIDIA Jetson Orin Nano. Separately there is 3,000 dollars for the best use of Tavily, 500 dollars each for twenty city winners, and 100 dollars plus NVIDIA merchandise for ten people who leave useful product feedback. City wins are for people who attended an offline event, so everyone there that night enters under Seoul. A single project can take one overall prize, or one track prize plus one bonus prize.

Submissions are a YouTube video no longer than three minutes, a public repository with an open-source license, a written description, and feedback on the Nebius and NVIDIA tools. Participants keep ownership of what they build. The physical AI track requires at least one minute of real hardware or a robot moving in the video. Simulation footage alone does not count, and much of what was said on stage that night ran straight into that condition.

Judging first screens for the basic requirements, then weighs technical implementation, design, potential impact, and quality of the idea equally. Judging runs December 1 to 15, and winners are announced around January 11. As of September 12, Devpost registrations stood at 3,589.

Nebius builds data centers alongside NVIDIA

The first session was led by Guan-ru Huang, a cloud solutions architect at Nebius. He is from Taiwan, has been with Nebius for over a year, and covers all of Asia including Japan and Korea.

The first thing he brought up was gifts: 100 dollars of credit for the Token Factory inference platform, and 1,000 free credits for Tavily, which Nebius recently acquired. If you are already using ChatGPT or a Google API, the same style of API endpoint lets you move straight over to open-source models. He also passed along what he called insider information: usually one team per city emerges as a standout, and Seoul is in that group.

When he asked who had heard of Nebius, not a single hand went up. Guan-ru laughed and said that was actually good. Nebius is already listed on Nasdaq, so when someone at an event tells him they are a small Nebius shareholder, he answers that he is doing his best to maximize shareholder value. Plenty of people know the name from its run in the US market, but what he wanted to stress was different: as a strategic NVIDIA partner, Nebius builds data centers to NVIDIA reference architectures, and it builds the hardware, the data centers, and the software stack itself, a vertically integrated structure. Large companies like Meta and Microsoft have validated it, and after the US and Europe it is expanding into Asia Pacific.

Putting up a layer diagram of the end-to-end AI platform, he said that if you understand every term on the slide you could work at Nebius, so come find him afterward, adding that they are hiring in Korea. He also introduced the Builder Program launching that day. Joining gets you 25 dollars of Token Factory credit, 25 dollars of Tavily credit, and a Nebius certification for one dollar. The official page also lists partner credits from LangSmith and Toloka. Everyone in the room already had their 100 dollars, he noted, so tell the friends who could not make it.

We hired the ML scientists and have no GPUs

Nebius describes itself as an AI-native cloud that combines the performance of a supercomputer with the flexibility of a hyperscaler. Guan-ru unpacked that sentence as two camps in the industry. Hyperscalers like AWS, GCP, and Azure built their data centers for general workloads rather than AI, while on the other side GPU rental shops lend you GPUs with no software stack. Nebius sits between them, building data centers for AI workloads from the start and layering cloud software on top.

The first thing customers look at when sourcing GPUs is time to market. For an AI startup, the worst case is not being able to get capacity.

The ML scientists panic. They have been hired, and there are no GPUs to train the model on.

Models today train for weeks or months, so if the cluster drops mid-run you start over. That makes reliability the next thing they check, along with whether developers fluent in Kubernetes, Terraform, CLI, and SDK can work the way they already work. A team might start with ten GPUs or a few hundred and need thousands as the business grows, so they also ask whether they can grow inside the same cluster and the same data center. Guan-ru returned repeatedly to the point that Nebius is not a bare-metal GPU rental but a full cloud stack.

He explained the hardware so non-engineers could follow. He called compute, storage, and networking, the three ingredients of a GPU cloud, the fat, carbohydrates, and protein of an infrastructure company. On a spec sheet, H100 and H200 are the older Hopper generation and B200 and B300 are the newer Blackwell. Each generation brings more GPU memory, and training a model with many parameters on older GPUs means splitting the model, which costs more maintenance and more developer time. That is why customers arrive having already decided which GPU they want.

A VM usually holds eight GPUs. GPUs inside a node are linked by NVLink and nodes are linked by InfiniBand, so if you need sixteen GPUs you connect two eight-GPU nodes. Even within Blackwell, the GB300 uses an NVIDIA-based CPU instead of the previous generation Intel CPU, so CPU and GPU sit more tightly together.

How you use the GPUs also comes in layers. The easiest is serverless: no VM or Kubernetes to manage, just push a container, run one-off training as Jobs and inference as Endpoints. For teams that want to handle infrastructure directly there is Soperator, an open-source operator that runs Slurm on Kubernetes. Today an ML scientist has been writing Slurm scripts since their PhD lab, while a DevOps engineer from AWS knows Terraform and Kubernetes but has never touched Slurm. Both of these people work at the same company, so Nebius gives them a Kubernetes cluster for DevOps and a Slurm cluster for the scientists.

Beyond that there are the widely used SkyPilot and Ray, plus managed Kubernetes. Storage splits into a shared file system that lets every node read the same data during distributed training, and S3-style object storage for data and training checkpoints.

Retraining the model on your own traffic

Next came Token Factory, the managed inference platform. Guan-ru asked whether anyone knew Artificial Analysis, the third-party site that compares speed and price across models and inference providers, and again nobody did. Nebius, he said, ranks in the top three there across models including DeepSeek and Kimi.

Developers today mostly use closed APIs or self-host. Closed APIs from OpenAI or Anthropic come with endpoints and documentation, so they are easiest at the start. But as a service grows, API costs swell and eat margin, and performance may not land where you hoped. Self-hosting on your own GPU cluster requires a dedicated DevOps team to run the inference cluster, and you may not be able to source as many GPUs as you want. The bigger problem Guan-ru pointed to is that neither approach gives you a chance to keep improving the model with your own traffic data.

Token Factory builds that loop into the product. You pick an open-source model and start inference, export the traffic data to see what customers are actually asking, then run post-training again. The resulting custom weights go up as a dedicated endpoint on the GPU type and region you choose. Guan-ru called it a virtuous cycle that keeps making the model better. Latency is handled with speculative decoding, cache-aware routing, and KV cache.

There are many inference platforms, but few own their AI cloud, so elsewhere you cannot choose GPU type or serving region. For enterprises that must meet security and compliance requirements, that choice matters. There are more than 60 open-source models to pick from and the list grows weekly, and going from PoC to production takes a few weeks.

No model has been trained on what happened two seconds ago

The last part was Tavily. Ask ChatGPT the current score of a game and the model has to fetch the answer from the web.

No model can be trained on that data for things that just happened two seconds ago.

The trouble is that web search results come back in formats like JSON that AI agents struggle to read. Tavily is a layer between the agent and the web. It offers search, which returns a list of relevant URLs; extract, which pulls the contents of a URL into something like markdown; and crawl, which follows relationships between URLs. The official site adds research for deep investigation and map for sweeping the URLs inside a site.

There were three problems to solve. The web, where everyone posts something new every couple of seconds, is not a static database, so accuracy is hard. Even with thousands of results, how relevant they are to the question is a separate problem. And letting an agent dig through hundreds of pages makes results slow enough that users leave.

Running through all three is freshness. If a web developer asks why the latest React upgrade is failing and you hand back documentation or a Stack Overflow question from two or three years ago, it is useless. For this Tavily runs a refresh scheduler. It combines four signals, how often a page changes, when it was last updated, how often it is searched, and how important the site is, into a single score that sets when to crawl that site next. Tavily is not its own model but a layer on top of models like ChatGPT, Mistral, and Anthropic, so it fits anywhere information has to come from the web, whether that is a coding agent, CRM data enrichment, or company research.

By the end of this talk, everyone is a roboticist

The second session was Dhruv Diddi, who leads the physical AI ecosystem at Nebius. He has worked at Google, Turo, and YouTube, and his area is the Physical AI Workbench, which bundles NVIDIA Cosmos, Isaac Sim, and Isaac GR00T. Dhruv opened by saying Seoul is a center of manufacturing, so he wanted to bring the GPUs, the robots, and this technology here. He first asked the room about backgrounds: traditional software was the largest group, some came from AI-native startups, and a handful were roboticists who already knew physical AI inside out.

My goal is that by the end of this talk everyone is a roboticist. Well, not literally, but at least comfortable enough with robotics and physical AI terms to hold a real conversation in them.

Physical AI means all AI deployed into the actual physical world. It could be an embedded system or making objects around the house smarter, but the highest-value form is robotics. The slide set digital AI beside physical AI. Digital AI outputs tokens on a screen, so a hallucination is just wrong text and the cost of an error is zero. Physical AI outputs motor torque wired straight into reality, so a hallucination becomes a physical collision and the cost of an error is fatal. A single instruction not to drop a glass has to end up as joint angles and torque.

In the physical world you cannot make mistakes. This is not about spilling a glass of water. If you are not precise, the consequences can be genuinely fatal.

He organized the terms we would hear all night into three groups: the two ways of training robots, imitation learning and reinforcement learning; evaluation, which you could call a test suite for robots; and foundation models. Where robotics people spend most of their time is simulation training and simulation evaluation, because before you put a model on a robot and destroy its motors you need to check whether it overheats or makes noise. One slide read that 95 percent in the lab becomes 50 failures a day at scale. Foundation models here also differ from LLMs, handling motor actions, trajectories, and sensor predictions, which is where NVIDIA Cosmos, GR00T, and Sonic sit.

The history slide was titled 76 years trapped in the abstract. It ran from Turing and rule-based AI in the 1950s, through the statistical era of learning from pixels from 1986 to AlexNet in 2012, into the generative era after the 2017 transformer, where AI reasons and creates in language but stays inside the screen. From 2023 onward, world models like Cosmos and action models like GR00T give AI degrees of freedom in reality. Dhruv said the next 75 years will belong to physical AI, adding that this is not his line but Jensen Huang's.

The three computers do not agree with each other

What Dhruv said you absolutely must understand about physical AI was the three-body problem. The factory that collects data and trains, the matrix that runs simulation, and the edge where the model actually lands are all different computers. The factory is a DGX-class GPU cluster of H100s or H200s; the matrix is an RTX-class server making digital twins and synthetic data with Omniverse, Isaac Lab, and Cosmos; the edge is an onboard computer like Jetson Thor that has to react in milliseconds. The reason robots are not everywhere yet is that the conditions of these three computers do not line up.

Sometimes data collection does not match simulation. Sometimes simulation does not match deployment. Sometimes nothing works, and sometimes everything works. The cool videos you see online are the everything-works case.

Every model in robotics divides into perception, reasoning, and action: a model that takes in sensory data, a model that reasons about what comes next from what it gathered, and a model that turns that reasoning into motor trajectories that move the robot. In the slide example, a VLM notices the glass is at the edge, a world model predicts that this trajectory will knock it off and break it, and a VLA corrects the grip at up to 1000Hz. Robots sound enormously complicated, Dhruv said, but they are not. You only need the frame: collect data, train, deploy the trained model.

He unpacked the model names with the same frame. Feed an LLM text and text comes out. Feed a VLM an image and text and it answers yes or no to whether there is a cat in this photo. Feed a VLA an image and text and out comes robot action.

A VLA does not answer with yes, no, or hello. It answers with what the motors should do, given this scene and this instruction.

A 200 dollar robot arm went around the room

Data for robots comes from three places. The most scalable is video, and he said many people are working on the problem because it would be wonderful if a robot could learn by watching YouTube. The catch is that phone video carries no force or joint information. Simulation lets a GPU cluster generate millions of episodes overnight, but the physics is an approximation and drifts from reality. Teleoperation, a person driving the robot directly, pairs actions and states perfectly but yields only five to fifty episodes an hour. The slide said the internet text is already spent, and that the data robots need, joint angles and forces interlocked with camera frames, does not exist at scale. At the end of this, Dhruv pulled out an SO101 robot arm for teleoperation.

This is an open-source robot arm. You can buy it online for about 200 dollars, and because it is 3D printed you can fix it yourself, break it yourself, and source the motors yourself. It is the most scalable robot you can get right now.

Move the leader arm and the white follower arm copies it exactly. You gather thousands of episodes this way through teleoperation and train on those trajectories. Dhruv sent the arm out into the audience, saying he hoped nobody would take it home, and it passed from hand to hand. Imitation learning and reinforcement learning are easy to separate by whether the robot interacts with an object.

Those robots you see online dancing and working out are all reinforcement learning, because they are not interacting with objects. To have a robot handle a cup of water or move a chip, you use imitation learning.

The principle on the slide was imitate first, reward later. Imitation learning, cloning human demonstrations, is fast and sample-efficient enough to reach an 80 percent success rate as a starting point within days, but it cannot beat the human. Then reinforcement learning, optimizing reward through trial and error in massively parallel simulation, grinds down the rare failures imitation learning never reaches. You have to randomize texture, friction, and lighting during training for the model to survive real-world variation.

All of this is complicated: the GPU setup, the training method, the know-how of handling robotics tooling, and the field is new enough that the tools are scattered. So Nebius bundled the workflow into a single open-source repository. It holds a training workflow that takes as few as five episodes and augments the rest into a model, a data augmentation workflow, and a robot model evaluation workflow that normally costs a lot of time and know-how. A CLI called npa connects your Nebius project and storage and pulls in LeRobot, Isaac Lab, Cosmos, GR00T, and FiftyOne, and any model deploys in three commands.

With the Workbench CLI you literally just tell an agent to do it. The agent sets up the GPU cluster and trains. We designed it to be agent-friendly so people who are not comfortable with a CLI, or have never done robotics, can jump in.

The prepared demo was a model trained to pick up a Nebius marker, but time ran short and it moved to after the talk. The example left on the slide was GR00T, pretrained on more than 20,000 hours of human video and robot demonstrations, fine-tuned about 2,000 steps for the SO101 and then asked to pick up a red block. A deployment slide he did not reach on stage carried this line: an inference delay of just 50 to 100 milliseconds can topple a bipedal robot, so intelligence has to be physically split between cloud and edge. A large VLA plans on a slow cycle while a small policy inside the robot holds balance and contact at 1000Hz.

The model builders, the field deployers, and the infrastructure providers sat in one row

From 6:15 I moderated the panel. Andy Lee, who leads APAC physical AI and the Inception program for startups and VCs at NVIDIA, joined Dhruv and Yu Been Park, CBO of Diden Robotics and also a Bloom member. Diden Robotics builds quadruped robots that climb the walls and ceilings of steel structures on magnetic feet to weld and inspect. They passed a Samsung Heavy Industries proof of concept and moved into delivery, and they are co-developing a welding robot for the inside of ship blocks with HD Korea Shipbuilding and Offshore Engineering. A company that builds models, a company that supplies the infrastructure to run them, and a company that actually bolts robots onto a shipyard were sitting in one row.

The first question was what physical AI they had touched with their own hands lately. Dhruv had given the same talk in Tokyo and was heading to Singapore next, and said that setting aside the big robots he owns about 35 open-source ones. Lately he has been working with OpenAI Astra, a model that people online say has shown emergent properties that make robot tasks easier. Andy had his 3D printer at home make life hacks that save Astra time, and recently bought a MicroDuck to poke around with. Yu Been Park's answer was short: his company builds robots, so the last one he touched was the spider robot he and a colleague carried in for the talk.

Asked to explain Cosmos and GR00T in plain language, Andy said this.

Cosmos is a world foundation model. It tries to understand the world we live in. It perceives, predicts, and generates video data so the model learns about its surroundings. GR00T is a robot foundation model. It turns what was learned about the world into actual behavior, into policy. So that when it sees a ball coming, it can raise a leg and catch it.

A perfect score in simulation still slips in the shipyard

Asked what is hardest about teaching a robot anything, Dhruv laughed that Andy had taken the easy question. The answer was dealing with the physical world. Simulation training has come a long way and GR00T and Cosmos are effectively state of the art. But as the previous generation of physical AI startups learned in autonomous driving, you score 100 in simulation, assume that is enough, deploy into reality, and only then realize you have to break the problem back down into autonomy levels like L1 and L2. Some things you only understand much later, including why they are complicated at all. Dhruv named deployment as the hardest stage, and said deployment gets smoother only when there is more of it.

I asked Yu Been Park how the robots actually work. He laughed that it was a very long question, then split the problem in two. One is navigation. A hull is a three-dimensional ferromagnetic surface, so the robot has to move vertically, horizontally, and along the ceiling, and getting from A to B uses model predictive control or reinforcement-learning-based control. The other is tool use. The industrial arm on the robot swaps between welding, angle grinding, shot blasting, and painting tools, and each tool controls differently, so they are working out how to control welding and coating on steel plate with a VLA model.

Then I asked, as the Nebius team had framed it, what breaks first when you run enough simulation and put it into reality.

When we first built the reinforcement learning simulator, we made a very ideal environment: flat floors, clean angles. But a real shipyard has disturbances. When you weld, small metal fragments scatter on the floor, and if the robot steps on one it slips and can lose control.

That is only the simplest example, and the variables that can go wrong are countless. Whether it is the NVIDIA simulator or MuJoCo, you can train the algorithm and test everything and it still breaks in reality. So to iterate quickly on the system you first need a real environment and a real robot, and you keep running it there.

Without the real world, you cannot build any kind of solutions to it.

I asked Andy whether the harder part of physical AI is the AI or the machine. The answer was firmly 50-50. AI models are improving exponentially, with GPT-6 Astra and the just-released Fable, but the moment you deploy into reality, variables you cannot account for pour in.

The lighting conditions are different, the environment is different, and something can go wrong. A few lights might be out. So the hardware has to catch up with the AI, and that is exactly the stage we are at.

Physical AI forces you to think more efficiently

Asked whether more compute makes a team better, Dhruv pointed the other way. Plenty of data, post-training, and good GPUs help of course, and that is thanks to the NVIDIA engine, a line that got a laugh. But the physical world forces specialization. Train on garbage data or data spread too wide and the weights never settle, so right now you do better by narrowing to one lighting condition, one robot, one environment. GR00T models also work well when you follow the guidance in the paper exactly.

In the end, physical AI forces you to think more efficiently. This is not the paradigm of dumping internet-scale data into a big model and making it work.

Physical AI models are in fact small, around two to seven billion parameters, because being multimodal lets you train them more densely. Where a text model only gains resolution by stacking text embeddings, a multimodal model has more space to fill, so he expects many clever ways of using GPUs to emerge. This was the point where the difference in approach between software and physical AI landed again.

Korea's advantage is that the factory is nearby

Here I asked something that was not on the question sheet. Software data looks similar anywhere, but physical AI is hard because you have to collect fine-grained data from the field, and Korea is known as a manufacturing and factory hub, so was there advice for the builders in the room? Dhruv said the advantage of being close to manufacturing sites and being able to walk into those spaces is clearly large. Korean startups are already using it, as the shipbuilding case shows, and he cited RLWRLD, a Korean company building foundation models for dexterous manipulation.

The biggest advantage comes from being able to control the manufacturing environment directly. That is why Korea is ahead on data collection and will be able to deploy models earlier than other countries.

I asked Andy why NVIDIA cares about Korea. Before this event I had no contact with NVIDIA and did not even know how large its Korea team was. Andy said Korea is highly strategic for NVIDIA. Manufacturing is strong across semiconductors and automotive, and the physical AI ecosystem is growing fast. By International Federation of Robotics figures, Korea leads the world with 1,220 industrial robots per 10,000 manufacturing workers, more than nine times the global average of 132. At the same time the population is shrinking and hiring for on-site work keeps getting harder, which makes demand from the manufacturing floor clear. Out of that come startups like RLWRLD and Motif Technologies, along with OpenGraph Labs, which collects tactile data from fingertips, and Space AI, which attaches physical properties like touch and force to objects visible only in video to make robot training data. Space AI is the only Korean company in the APAC AI accelerator run jointly by NVIDIA and a Taiwanese accelerator. Andy said the region is full of hidden gems.

When Andy asked for a show of hands on who had heard of NVIDIA Inception, not many went up. Inception is a program for startups and VCs, and any company incorporated within the last ten years can join for free. It brings the chance to learn and build with NVIDIA products like Nemotron, Cosmos, OpenUSD, and Omniverse, Nebius cloud credits, and introductions to US and Korean VCs for fundraising. With the NVIDIA Korea team here able to connect through to the US and Europe, that is a real advantage for Korean startups.

The Nebius site also carries a call for the Nebius Physical AI Awards for physical AI startups. It is open to teams that have raised over a million dollars and are running a product in the real world, with applications until October 21. Category winners receive 150,000 dollars in Nebius compute credits and present at the San Francisco summit in December.

Do the hardware

I asked Yu Been Park what had been hard looking back on founding the company, and what he would do better if he did it again. There was no hesitation in the answer.

In physical AI, you have to do the hardware. If you use external hardware there is a clear limit to how far you can optimize the system as a product.

If you are a genuinely strong software developer, he said, find a hardware co-founder. When you are only optimizing the software layer, it is far too easy for a customer to pull your software out and swap in another vendor. If you take Chinese hardware off the shelf as is, the gap between you and a team that is better at software disappears, and as language models improve, software gets easier and easier to replace.

How do you stay attached to the customer? It is simple. Do both hardware and software, integrate it all vertically, and sell it as a package.

The last question was what single piece of advice they would give someone starting to build today. Dhruv said start with the resources already out there, including the Builder Program. With the right tools you can make magic, and you do not have to lean on AI for everything; you can mix MCP into it. His yardstick was ROI: focus on what is impressive and high-impact for your team, and on what your team brings to this competition.

I would recommend to start small. Rome was not built in a day. What you build tonight may not be the best product and may not give you the result you hoped for. Do not aim for the best model in a day. Set small goals in small steps, measure whether you actually hit them, and move from there.

The mic went last to Yu Been Park.

I am not a hackathon winner either, so I cannot give great advice. As I said earlier, software is fairly easy to build these days. So how about differentiating with hardware? That is what I think.

The most valuable data is collected while deploying

After the panel the hackathon began. For latecomers, the Nebius team walked through claiming credits and what to build again over Discord. Credits came in two kinds, cloud credits and Token Factory credits, 100 dollars each, claimed through a QR posted in Discord, though compute credits went only to people who had registered on Luma. The reason given was that they wanted people to use compute responsibly.

The first instruction was to form teams. Any team that formed and decided what to build could come up at 8 p.m. and introduce it for a Nebius t-shirt, first come first served, with the condition that you could not just say anything.

Meanwhile Dhruv continued the data conversation with people who came up with questions. Besides the teleoperation he had just shown, there is an approach where you put a gripper-shaped device in a person's hand and have them work. That is easier than transferring human hand motion into a gripper and then into a robot, so gripper data is cheaper to collect. Collecting robot to robot directly is worth more, and the most valuable of all is data collected while deploying the model. If you have a model automating some task, however clumsily, you let the workers watch and correct it, and the data that accumulates that way is called DAgger data. Dhruv called it the most effective and highest-ROI data there is.

There was one more thing Dhruv said after coming off stage. This Christmas, children might be unwrapping robots under the tree. San Francisco already feels that way, he said. Considering the robot arm that went around the room that night is a 200 dollar open-source design you print yourself, it is not a distant story. Hardware, too, is becoming something anyone can stamp out.

Another gap Dhruv named was embodiment. You cannot take what you trained on this robot and use it on another one. There is still no way to make cross-embodiment work between different robots, and Dhruv sees exactly that gap as the opening for innovation. Unlike AI, the physical side is still empty land with far too much to do, which is why he says the world needs thousands more roboticists.

We said no demos and no pitching, and more than ten teams took the mic

At 8 p.m. people who had formed teams and settled on ideas came up one by one. The Luma page had said no demos and no pitching, but once we said going up alone was fine, more than ten teams took the mic. Dhruv gave feedback on the spot for most of them. Presenters are described here only by what they do.

The first presenter said he wanted opinions directly, and that since the idea was not settled it was fine to attack it or tear it down. He does not think building one more agent is a differentiator, and wants to solve the social problems that appear when the economic actor is an AI agent rather than a person. For example, one hotel booking agent queries prices at a million hotels worldwide, places bookings, deliberately holds them for 23 hours, and cancels them all. The hotels have to service useless requests and hold rooms, and tens of thousands of such agents would paralyze the booking system and the economy around it.

A buyer agent costs zero. Anyone can build one. But a seller agent has to look up prices, negotiate, and actually take the booking.

What he wants to build is a commitment gateway that helps the seller side so requests do not become free. It tells the seller whether this buyer is the kind that queries a million listings to book one, or the kind that looks at five or twenty and books one. Without hard-capping call volume, it separates real buyers from real sellers to help an agent-to-agent economy actually function. Dhruv said it is a perfectly valid problem, and that a lot of ideas about agent-to-agent structures are surfacing right now.

Features that are easy for AI to use end up easy for people to use

A team building a Kubernetes-based GPU management platform started in English and switched to Korean. They said we might laugh at them for doing something similar to Nebius, but they do not provide infrastructure themselves; they build software that sets it up and helps people use GPU servers well. GPU infrastructure has always been hard for non-engineers, so they had focused on UI that is easy for people, and then using AI changed their thinking. Just as AI now drafts work that used to require a person to research and build slides, infrastructure setup could work the same way. For example, when a university professor gives 30 students a practice environment, the infrastructure setup usually handed to a teaching assistant could be done by AI.

I think features that are easy for AI to use will end up being features that are easy for people to use.

Dhruv said GPU orchestration is not something to view as competition. The field is still early and there is a lot to optimize, especially at the level of interaction UI, so think of it as growing the pie rather than competing, and definitely build it.

Someone who runs a YouTube channel said none of the editing tools he wanted existed, so he built his own. Using it on his own channel, he got feedback that a lot of people around him wanted it too, and learned that many people who want to publish content cannot because editing is hard. His goal is to help those people so more good content reaches the world.

Someone who wants to build a personal AI used Iron Man's Jarvis as the reference. The most basic piece is notifications. The trigger was something that happened recently: he missed the deadline for a residence registration survey, so someone is about to come to his home, and if AI had seen the address and told him in advance it would already be done. So he sketched a system that finds the deadlines you have to keep, summarizes them, and points you at the right site. Public offices push out information, AI gathers it and alerts whoever needs it, and events posted on Luma that match your interests arrive as notifications you can register for immediately, helping both the side announcing and the side receiving.

Not the duckling but the egg the duck came from

Someone who introduced himself as a maker came up with a sample he had built. He has twelve 3D printers in his office, and ten years ago he built ten 3D printers himself through the RepRap project. He worked as an English instructor, became a public servant and ran a makerspace for three years, then quit when it stopped being fun and started a company. His office holds thousands of parts, and when he thinks of something he wants to make he has a prototype in two hours. He was honest about why he makes so many educational products.

Between a product and a prototype, education is the only place you can actually make money.

His proposal was an egg. He drew a lot of inspiration from MicroDuck, the open-source duck robot from Hugging Face and Pollen Robotics, and said he would build not the duckling but the egg the duck came from. Using the reaction torque approach that controls satellite attitude, an egg-shaped robot can roll around on its own. Add lighting and a face and it would be cute, he said, sketching a robot under 200 dollars, small and charming enough that you would want one, wandering around the house. He does not know yet what the egg is for and asked for inspiration, and since he does not know GPUs well he was looking for a partner to handle the technology.

Dhruv said he loves open-source ideas made with 3D printing and MicroDuck in particular, and that the egg is a good form factor. What you always want from a robot is interaction, so he suggested adding actuation to an egg you can talk to, something like a small Reachy Mini. Voice is good, but it is better if the egg can knock to make a sound or control other objects.

I think the real value of a robot being present in reality is that it can move something.

It started with adding an arm and went as far as an egg with only legs walking around being fun too. The egg that nobody knew what to do with started growing arms and legs right there.

Pressing a button is a high-difficulty task for a robot

One team wanted to use the task of pressing a button to work on differences between robots. Cross lighting, so it works when the lighting changes; cross environment, so it works when the environment changes; and how to map the physics of one motor onto another came up as their challenges. Dhruv's first piece of advice was not to invent a new task. Tasks already in the benchmarks have models optimized against them, so look at what the models are best at and build a fine-tuned version for a different robot.

He went on about why pressing a button is harder than it sounds. There is timing in the moment of pressing and the moment of releasing, and press too hard and the button breaks, release too early and it does not register.

Most of these models judge by vision, not touch. They do not know whether they have made contact, only whether the color of the robot arm is near the button or far from it.

So in robotics, pressing a button takes a lot of dexterity. Instead he suggested starting with something easy like picking things up, and with soft objects unlikely to break, like a t-shirt, a squishy toy, or socks. The presenter asked back: what about putting something like a QR code on the button so it knows the orientation and direction? Dhruv said now the computer vision tricks are coming out, and that a QR works, or you can specify where the robot should go with a box, so be a little clever and make it work reliably.

A researcher at a company that helps write radiology reports asked about training cost. Training at scale takes a lot of resources, the data itself is scarce, and data across many dimensions needs to be accurate but often is not. Starting from the question of whether you have to pay a large cost to train every time, they want to research training quickly with small language models or multimodal models and delivering something tailored to a relatively small user group. Their question was what to focus on when the areas you can automate in research keep growing.

Dhruv said the opportunity is in data augmentation. Physical AI especially is short on data, so if you have 100 episodes you want 1,000. For a language model you generate synonym sentences that teach the same concept in slightly different words; for vision you scale an image up 10 percent, down 10 percent, crop it, or tilt it five to ten degrees, multiply the data tenfold, and fine-tune. It is already used in practice, and in research too he said to focus on the highest-ROI work. For an open-source library to do this with, he pointed to Oumi.

I do not have feedback for you, I am a fan

A neurosurgeon is building an autopilot for spinal surgery full time. Stating a goal of conquering every surgery in the world with robotics, he said that just as data was the bottleneck for digital AI, physical AI models have to collect as much data as possible. The problem is that data is hard to get in an operating room. So he is collaborating with a company that makes spine dummy models to collect the positioning surgeons already do in the operating room, and he has already built a tip with a tracking system on it. He was looking for someone with a computer vision background to join.

I do not have feedback for you. I am a fan, actually.

Dhruv said he had picked an enormously hard problem both technically and in terms of impact, and admired that he had already made progress with a device built with his own hands. Ways to improve it by adding more sensors, he said, they had already discussed earlier.

A team building an AI autobiography service for older adults is running it as a program at several welfare institutions. Meeting older adults, introducing AI, and teaching them to use a phone, they find it harder for people to follow than expected, and even where an app like ChatGPT is installed almost nobody has actually used it. So they plan to train a Korean-specialized model into a small local SLM and build it into a compact device older adults talk with in daily life. Older adults are not going to code or solve math problems, so they judge that focusing on Korean makes a cheap, small model sufficient. They named the product AI, a name you call the way you would call out to a baby in Korean.

I think the big opportunity in AI is access to intelligence. Older adults are exactly the people who need intelligence they can use without buttons or typing.

Dhruv said he loves projects with large social impact. There is a big market for Korean-specialized models but not enough Korean data, so releasing the data they gather as open source would be good. Deployed on-device it can help an older adult 24 hours a day offline without relying on the internet, and post-training can happen in the cloud.

You need a foreign language to make friends, and friends to learn a foreign language

As the presentations were wrapping up, someone who had missed the chance earlier took the mic saying we did not have to listen. What he is building is a voice messenger without a language barrier. He wants to get close to people from other countries but finds the language barrier too hard to cross, and there is a chicken-and-egg problem in it.

I want foreign friends, but I need a foreign language to make friends, and I want to learn a foreign language, but I need foreign friends to learn it.

There are many language exchange services but none has really solved this, and other services that connect you with people abroad end in giving up at the language barrier. So the proposal was to dissolve the foreign language itself and let people focus only on hanging out with friends from other countries. Looking into why existing services could not do real-time interpretation, he found that older AI could not meet the demanding requirements of a social situation, and there were UI and UX problems too. From this year all of that became solvable thanks to AI progress, and he realized this is a business that only just became possible.

He spoke Korean into the interpretation app he built and real-time English and Japanese translation flowed out. Pairing it with HelloTalk voice rooms got a good response, and a fair number of users now enter voice rooms with this translator running. Many use it offline or for work, and some have landed jobs with it. There are about seven usage scenarios that need interpretation and he plans to take them one at a time, starting by drawing in K-friends users. Since no service fills both real-time interpretation and global socializing, the plan is to start from interpretation and move toward an AI-native messenger and social network. He was also looking for teammates.

We decided to do what only offline can do with AI

Something from a conversation with the Nebius team after the event has stayed with me. If you are an offline community, hold on to what only offline can do. Physical AI fits that condition exactly. Anything that ends inside a screen works from anywhere, but building a robot, touching it, and breaking it takes being in the same room.

Inside the team we also tried applying what Yu Been Park said to Bloom. Software scales easily but is now far too easy to replace; hardware is hard to scale but hard to replace. What counts as hardware at Bloom is, in the end, people. AI has become something anyone can use and control, but people still do not bend to your will, which is why rooms where people gather hold their value. That is also where the line about coming to Bloom to see people instead of a screen comes from.

The threshold has dropped a lot too. The SO101 that went around the room is a 200 dollar open-source arm, and MicroDuck from Pollen Robotics, acquired by Hugging Face, is taking pre-orders at 399 dollars. It is a 25 centimeter, 800 gram bipedal robot with 15 motors, and its SDK, simulation, and reinforcement learning stack are all open under Apache 2.0, so you can retrain all seven of its built-in motions. Developer response has been good enough that pre-orders are backed up, and ordering now means waiting until the end of the year. It is not only robots. A 359 dollar Linux board built for keeping agents running is also taking pre-orders.

This is why Dhruv talking about Christmas did not sound like a metaphor about the distant future. So inside the team we decided to try two things. One is a small offline session where participants assemble and train open-source robots like the SO101 or MicroDuck themselves; the other is a regular physical AI meetup connected to Nebius and NVIDIA. We also talked about Bloom possibly being the one to bring robots like these into Korea. The point is doing, with our own hands, what only offline can do with AI.

These two will not be the end. As the team building AI for older adults showed that night, if AI becomes necessary regardless of generation, then Bloom has to step outside the circle of people who do AI for a living, someone on the team said. I am not sure yet whether that is right. What did become certain that night is that some things only happen outside the screen.

I do not know how many of the teams that took the mic will actually submit by the 2 a.m. deadline on October 31. If today's robot is something that scores 100 in simulation and still slips on a single metal fragment, seven weeks may be short. Still, I am quite curious what the maker who brought his own sample and the doctor who showed a spine dummy model will bring by then.

FAQ

What is the Nebius and NVIDIA Global AI Hackathon? A twenty-city global hackathon run online, with submissions due 10 a.m. Pacific on October 30. Builds must run on Nebius Token Factory or AI Cloud and use at least one NVIDIA open-source model, across four tracks including physical AI, with more than 50,000 dollars in prizes.

Why is physical AI harder than digital AI? Digital AI outputs tokens, so an error costs nothing. Physical AI outputs motor torque, so a hallucination becomes a collision. On top of that, the three computers involved, training, simulation, and edge, do not line up with each other, and a perfect simulation score still breaks on contact with reality.

What is Korea's advantage in physical AI? Proximity to and control over manufacturing environments, which puts Korea ahead on data collection. Korea also leads the world with 1,220 industrial robots per 10,000 manufacturing workers, more than nine times the global average.

Scenes from the night


Join Bloom

Bloom builds offline rooms where people and technology meet. We run them in Seoul, and now beyond it.

Read more

ํ”ผ์ง€์ปฌ AI๋Š” ์˜ฌํ•ด ํฌ๋ฆฌ์Šค๋งˆ์Šค ์„ ๋ฌผ์ด ๋  ๊ฒ๋‹ˆ๋‹ค

ํ”ผ์ง€์ปฌ AI๋Š” ์˜ฌํ•ด ํฌ๋ฆฌ์Šค๋งˆ์Šค ์„ ๋ฌผ์ด ๋  ๊ฒ๋‹ˆ๋‹ค

๐Ÿ“… 2026๋…„ 9์›” 11์ผ ๐Ÿค Bloom ร— Nebius ร— NVIDIA ๐ŸŽค Guan-ru Huang, Nebius ํด๋ผ์šฐ๋“œ ์†”๋ฃจ์…˜ ์•„ํ‚คํ…ํŠธ ยท Dhruv Diddi, Nebius ํ”ผ์ง€์ปฌ AI ยท Andy Lee, NVIDIA APAC ํ”ผ์ง€์ปฌ AI & Inception ยท ๋ฐ•์œ ๋นˆ, ๋””๋“ ๋กœ๋ณดํ‹ฑ์Šค CBO ๐Ÿ“ ํ•ฉ์ • ์šฐ๋ฌผ ๐ŸŽŸ๏ธ ํ–‰์‚ฌ ํŽ˜์ด์ง€ ๋ณด๊ธฐ Bloom๊ณผ ํ•จ๊ป˜ ์„œ์šธ์—์„œ ๋ฉ‹์ง„ ํ–‰์‚ฌ๋ฅผ ๋งŒ๋“ค์–ด์ค€ Nebius์™€ NVIDIA์— ๋ฐ•์ˆ˜๋ฅผ ๋ณด๋ƒ…๋‹ˆ๋‹ค. ๊ธ€๋กœ๋ฒŒ AI ํ•ด์ปคํ†ค์ด ํ•œ๊ตญ์—์„œ ๋ง‰ ์‹œ์ž‘๋์Šต๋‹ˆ๋‹ค.

By Bloom
์—ฌ๋Ÿฌ๋ถ„์˜ AI ์„œ๋น„์Šค, ์ •๋ง ์•ˆ์ „ํ• ๊นŒ์š”?

์—ฌ๋Ÿฌ๋ถ„์˜ AI ์„œ๋น„์Šค, ์ •๋ง ์•ˆ์ „ํ• ๊นŒ์š”?

๐Ÿ“… 2026๋…„ 9์›” 10์ผ ๐Ÿค Bloom ร— Datadog ร— LG CNS ร— ๋””์บ ํ”„ ๐ŸŽค ์ „์ฐฝ์›, LG CNS AI์„ผํ„ฐ Lead Researcher ยท ์†์šฐ๋‘, Datadog Sales Engineer ๐ŸŽŸ๏ธ ํ–‰์‚ฌ ํŽ˜์ด์ง€ ๋ณด๊ธฐ ์š”์ฆ˜ AI๋กœ ๋ฐ๋ชจ ํ•˜๋‚˜ ๋งŒ๋“œ๋Š” ๊ฑด ๋ˆ„๊ตฌ๋‚˜ ํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. ํ•˜์ง€๋งŒ ์กฐ์ง ์ฐจ์›์—์„œ AI๋ฅผ ์‹ค์ œ ์„œ๋น„์Šค๋กœ ์•ˆ์ •์ ์œผ๋กœ ๊ตด๋ฆฌ๋Š” ๊ฑด ์™„์ „ํžˆ ๋‹ค๋ฅธ ์ฐจ์›์˜ ์ด์•ผ๊ธฐ์ž…๋‹ˆ๋‹ค. ์™œ ์ด๋ ‡๊ฒŒ ์–ด๋ ค์šธ๊นŒ์š”? ๋น„์šฉ, ๋ณด์•ˆ, ํ˜‘์—…

By Bloom
Stop formatting proposals. Start winning them. Try Contrl Free Join Beta