H3-metal – Native MiniMax-H3 inference for Apple Silicon

https://opengraph.githubassets.com/5de2b0d9b61e80ca8935c2ae9fcc564be437066d396ba74b091f478c9e5e5875/antirez/h3.c
Native minimax-h3 inference for apple silicon. specialized 256-thread kernel gathers and quantizes each h3 row directly into the projection's row-major int8 buffer, eliminating the intervening full-width bf16 transpose without changing any output byte. cross-m5 measurements improved complete 512 forwards by about 0.2-0.8% avoiding a second device-memory read when

Show HN: Mcptoon – Token-efficient MCP CLI client

https://opengraph.githubassets.com/b9fc90e66635ec538b3ae53d725424422c01d05aa039a88149ca79d366cba487/activeing123/mcptoon
Mcptoon is a cli client that connects to any mcp server (stdio or http) it outputs toon (token-optimized object notation) instead of json. works with every ai agent — claude code, codex, opencode, cursor, catpaw, anything that runs shell commands. configure the servers, tools, token savings once - every agent shares the

Chicken Scheme 6.0

- fix warnings about zero-sized memset in build system. fix'memory-statistics' by returning semi-space bytes like the documentation says. increase the "binary compatibility version to 9' a number of bugs have been fixed in the latest version of libchicken, including the following: r4rs no longer exports eval and srfi-

As AI eats the web, the internet’s collective memory is disappearing

LFM2.5 2.6B model competitive with 4x larger models

https://cdn-uploads.huggingface.co/production/uploads/644249b08443bce4c9890a0f/NAfUL2725b6KDUGf1KOd9.png
Lfm2.5-2.6b is part of a family of hybrid models designed for on-device deployment. it builds on the lfm2 architecture, with an 128k context window and agentic post-training - the fastest model in its size class, reaching almost 15k output tokens per second at high concurrency, roughly 1.3b token per day on samsung h100s. the model is optimized for cpu inference

Show HN: Scroll through all 43252003274489856000 Rubik's Cube states

https://everycube.alen.is/opengraph-image?e65a545ab414fc11
Scroll through all 43,252,003,274,489,856,000 reachable Rubik's Cube permutations.

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

https://cactuscompute.com/_next/image?url=%2Fneedle2%2Farchitecture.png&w=3840&q=75
Cactus needle is a 2bit model that fits 45m parameters into 14mb. the model is trained on 115b tokens and post-trained on 38b engrams - jim snyder & joe wilson: the performance never lets us down. he says all numbers are measured end-to-end through the shipped c++ engine in its production configuration. "the tables answer one

The “mechanical miracle” that ruined Mark Twain’s life

Recycle – Floppydisks

https://static.wixstatic.com/media/f3bcf8_f5c6d3965bcf44479f00eec297e5d51b~mv2.jpg/v1/crop/x_3,y_0,w_114,h_120/fill/w_111,h_116,al_c,q_80,usm_0.66_1.00_0.01,enc_avif,quality_auto/f3bcf8_f5c6d3965bcf44479f00eec297e5d51b~mv2.jpg
We buy new disks in sealed undamaged packs. Quantities of 100 or more. If disks are outside their original packaging, we cannot treat them as new. For new disks, send us a picture of your disks, and call (800) 397-7890 for a quote.

Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models

The following information can help our support team to resolve this issue.

The UK's war on anonymity has come to America

https://www.effort.news/images/uk-lobby/ccdh-fara-letter-screenshot.png
Effort investigation identifies coordinated operation to influence american lawmakers. five foreign non-government organizations and their us affiliates are involved, authors say syria khan and edward mccaffery report reveals. they are attempting to pass a patchwork of digital id laws in 21 states and the us congress - they say. cnn's john defterios and david

Stowaway – Take the window seat on any plane or satellite overhead

https://stowaway.live/og-image.jpg
The sky here is the real one outside — your location, your weather, this minute's light. click one and the camera follows it; stow away and it takes a seat on board, over real terrain. all of that is drawn in real time, which needs javascript and webgl 2. switching javascript on for this site should do it. if you want to use the site in mobile, go to http://www.satellite

Sonic Pi v5

Rust SIMD on the GPU

https://www.vectorware.com/_next/image/?url=%2F_next%2Fstatic%2Fmedia%2Fvectorware_logo.6d5f5210.png&w=640&q=75&dpl=dpl_4ZmndNCZY2uSnEayhwbSuemEqYJP
Vectorware is building the first gpu-native software company. we are mapping rust's portable simd onto the gpu's native execution model. the ir itself is architecture-agnostic, so code and libraries can be used on gpus without rewrite. if you are interested in learning more, contact us at [email protected] - e-mail: adam@v

What's the best programming language for coding agents?

A post suggests that languages that represent things more concisely are more token efficient. the median time to solve the tasks is in the hundreds of seconds - sam taylor-mcgraw-smith et al. he says this is not always obvious how much the issue generalizes across tasks or across setups. for these kinds of data-y projects, llms massively reduce the amount of effort it takes to get

Confessions of a Long-Distance Sailor

https://arachnoid.com/lutusp/bt1_sm.jpg
In 1988, i sailed away from the west coast of the united states in a 31-foot sailboat. this is my story — my account of... around the world solo sail - for 3 1/2 years, from fearing shore, to wanting never to touch land again. read my book, sail with me! " ''i'm not sure if it's possible,' he says. but it is

How Claude marks AI-generated content

https://downloads.intercomcdn.com/i/o/487548/17213f6a445c8e6e874b1f4b/fad85208982e639d11b9108df895a293.png
Claude models launched in the eu on or after august 2, 2026 will support machine-readable marking at launch. embedded watermarks and digitally signed provenance metadata will apply to all generated text - and files if supported by the platform ad hoc or the underlying file type. lack of detecting mark doesn’t mean the content was not ai-generated or processed – it just indicates that it may ...

Publishing Schematics Before “Open Source” Was a Word

This website is using a security service to protect itself from online attacks. the action you just performed triggered the security solution. you can email the site owner to let them know you were blocked. email them with the cloudflare ray id found at the bottom of this page. if you have any questions, please contact [email protected] - we'll be happy to help. click here for ...

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

https://lookaside.fbsbx.com/elementpath/media/?media_id=1786825739136571&version=1786319248
Meta Superintelligence Labs introduces Muse Glimmer, a 30-billion-parameter model optimized for local agent workflows, enabling use anywhere, anytime. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks, with optimized integrations and developer documentation available for building and running agents.

Faster floating point math with Rust's new API

https://pythonspeed.com/assets/titles/faster-float-math-rust.png
Floating point math is often slower than integer math because the compiler is being conservative. starting in version 1.98, rust will allow telling compiler it can optimize your code further, but with extra control so that you can still write numeric algorithms with minimal rounding errors. if you want to sum an a large number of floating point numbers, add numbers with algebraic operators ...

Squeak 6.1

https://squeak.org/static/img/balloon-og.png
Squeak 6.1 release notes highlight major changes, including a new tree browser, improved kernel infrastructure, and enhanced toolset and user interface.

Tail-call optimization in C is relatively recent (2025)

The C calling convention has historically not supported tail calls, but GCC now supports them with a separate calling convention. This allows for improved performance in interpreters like Gforth and a toy Forth interpreter.

Hyperspace

https://hypercritical.co/hyperspace/images/hyperspace-icon-256.png?v2
Hyperspace finds files with identical contents and then reclaims the disk space taken by all but one of the identical files. if a file is marked as “busy,” it will be skipped, or it could have been set moments ago by an app that is using the file - but it won't do so unless it's owned by the current user. certain directories are entirely ignored by hyperspace, so an application might ...

Learning more about Claude's mathematical capabilities

https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F406d8496be38ea8a5459c66eaa7c5c7edab4a62f-4107x2310.png&w=3840&q=75
Claude has improved on a longstanding lower bound for the fraction of zeros that satisfy the riemann hypothesis. the result draws heavily on techniques that allow the hypothesis to be true without it, john sutter says - and he says it's not surprising that it did! if you're interested, you can find the paper here, or read the full explanation here. d.j. suriajaya: the results

Choral: Choreographic Programming for Java

Choral is an object-oriented language with a twist: objects have types of the form t@(r1,..., rn) roles are the roles that collaboratively implement the object. in our first use cases, chorals reduced the lines of code that we had to write to implement choreographies from 50% up to almost 400% - sam taylor, edward mcgraw-hill, andrew

Show HN: Ante, a coding agent in a single binary that runs offline

https://raw.githubusercontent.com/AntigmaLabs/ante/main/docs-site/static/assets/cookbook/providers.gif
Ante is a self-contained coding agent that lives in your terminal. it can also be the optimized core for building your own harness and high-performing assistants. ante is fast, lightweight, and the only terminal agent with native local inference built in. read more in our philosophy and agent organization patterns. to get started, download the alpha version of antigma.ai/eval.xml - it's available now

Humanising LLM Outputs Is Dumb

https://i.ytimg.com/vi/ExaNNAIysks/maxresdefault.jpg
Agents are increasingly patching this at the prompt layer, something that belongs further down the stack, says sam taylor. agents whose native language is precise, machine-facing state, with the warm, concise, human version generated only a few layers down, he says. if we want to make agents more efficient, we need to rethink the way they communicate with each other, and how they interact ...

Exploiting System Management Mode with a very long interrupt

https://opengraph.githubassets.com/f2afe945b79d47d3803682870ab4fa09e69e55f1dff3a9f323471c46fccbb986/xoreaxeaxeax/smiiiiiiiiiiiiiiii
Exploiting system management mode with a very very long interrupt. to break this, all we need is someone too busy to notice they're supposed to join smm - at this point, core 1 can attack core 0 while the other is out of it. there are 100+ smm toctou cves out there: an mmap handler checks and then uses it; all you need for exploitation is to rewrite

Mars Bar from 1991 found – and it's 20g bigger than today's

https://ichef.bbci.co.uk/news/480/cpsprodpb/79f0/live/51bb04a0-94ad-11f1-afe9-fb1a837ec5d9.jpg.webp
A 35-year-old mars bar has been found during house clearance. the chocolate bar was found alongside 'today's' 40g mars bars in scunthorpe, lancs. owner of cleaning service pocket rockets said it was "as big as my hand' she said she might re-sell the bar to fund upcoming uk tour - but is not sure what she will do with it

Parametron: 50s Japanese computer that uses neither transistors nor vacuum tubes