DeepSeek V4 Flash on a Single AMD MI300X

https://opengraph.githubassets.com/d869ff3c36cbcc543201134baed899936dfa6f93cdc17b3e9b24a1402cd31ef4/ryanzhou/deepseek-v4-flash-mi300x
This repository contains a production configuration for running DeepSeek-V4-Flash-0731 on AMD MI300X with fixes for FP8 format, MoE routing, and other issues. It achieves 6,988-7,019 tok/s on fresh prompts with a 2,048-token scheduler budget.

Ray Bradbury's "There Will Come Soft Rains" is set today (2026-08-04)

https://d2se8qvny5nxid.cloudfront.net/eyJidWNrZXQiOiJzaG9ydC1zdG9yaWVzLXdlYiIsImtleSI6ImltYWdlcy9wcm9maWxlLWltYWdlcy9pWW5VUmZQNUpZeW1vaElCR25RMDFUU1dCMERqZi5qcGciLCJlZGl0cyI6eyJqcGVnIjp7InF1YWxpdHkiOjgyLCJwcm9ncmVzc2l2ZSI6dHJ1ZSwidHJlbGxpc1F1YW50aXNhdGlvbiI6dHJ1ZSwib3ZlcnNob290RGVyaW5naW5nIjp0cnVlLCJvcHRpbWl6ZVNjYW5zIjp0cnVlfSwicmVzaXplIjp7IndpZHRoIjoyNTAsImhlaWdodCI6MjUwLCJmaXQiOiJjb3ZlciJ9fX0=
A house stood alone in a post-apocalyptic city, its automated systems still functioning. The house's voices sang and announced the time, but chaos erupted when a fire broke out, destroying the house and its inhabitants.

Xbox goes down. You can't play games you own on disc

An Xbox outage blocked users from playing both digital and disc-based games. The author prefers PC gaming because it allows for maintaining access to games despite digital issues.

LLMs reward expertise

https://www.seangoedecke.com/og-image.jpg
Using large language models like LLMs requires domain knowledge to get the most out of them. A skilled prompter like Terence Tao can steer the model with expertise in the domain, but this skill is not easily replicable by following tips.

Buckminster Fuller: everything I know

https://i0.wp.com/www.bfi.org/wp-content/uploads/2018/11/Buckyarticle-940.gif?fit=940%2C470&ssl=1
Buckminster Fuller's 42-hour lectures cover his inventions, discoveries, and design approach, spanning topics from architecture to economics. The lectures are now available as a transcript, inviting public participation in editing and refining the content.

Harness Engineering for Self-Improvement

https://lilianweng.github.io/posts/2026-07-04-harness/openai-agent-loop.png
Harness engineering is a crucial component of recursive self-improvement (RSI) in AI, enabling models to improve their own performance and adapt to new tasks through workflow design, context management, and optimization. Recent research has focused on developing automated design of agentic systems, meta-harnesses, and evolutionary search methods to improve harness engineering, with ...

AI-Generated Images Discourage Me from Reading Your Blog

https://nelson.cloud/ai-generated-images-discourage-me-from-reading-your-blog/og.png
You dislike AI-generated images in blogs, especially in indie blogs, and prefer authentic content from real humans. You value the unique perspective of individual bloggers over potentially AI-generated content.

Show HN: Fine-tune an 8B model on a 4 GB laptop GPU

https://raw.githubusercontent.com/MakazhanAlpamys/Soup/main/soup.png
Soup is a tool for fine-tuning large language models with a simple workflow, using one config and one command. It supports various models and training tasks, including LoRA, quantization, and batching.

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

https://opengraph.githubassets.com/33c2323be10c8fc55b1ca1698918168f8c99940b20885c0e596079d869780dac/leonickson1/Swiftlet
Swiftlet is a Swift + Metal runtime for Qwen3-Next and Qwen3.5/3.6 models, allowing them to run on iPhone with 2.5 GB of RAM. It streams weights from storage and activates only 3B parameters per token.

Roame (YC S23) Is Hiring Lead Engineer

https://bookface-images.s3.us-west-2.amazonaws.com/logos/a16a93b3a2821d2403e09535a2d3d3df3024fc8f.png?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=ASIAQC4NIECAATCC3II2%2F20260804%2Fus-west-2%2Fs3%2Faws4_request&X-Amz-Date=20260804T120518Z&X-Amz-Expires=3600&X-Amz-Security-Token=IQoJb3JpZ2luX2VjEEEaCXVzLXdlc3QtMiJIMEYCIQDYE19ElNBuEuktkmQRc5jVmRONVPW4FixBs45R52dpnwIhAIcWtXpXfCZu11ybavh4F9yQ3sL7AHkenLT3sJaWS7zYKuUDCAoQABoMMDA2MjAxODExMDcyIgwDo%2FHDq5peCbIvrjUqwgO4jBQBCpbd2OMtQdQ9AqmOXjfCd0rUabnLrDvecn6FsOYdxCV06hQr0oMTEpZWJPBMm7dhP9T9NSfqppxm8aoqXbmX8wEjs0mXcvTF01KtliztqQ0SZVOGtQRmdFpWl5W1mGL2qAnIrOSxFcMWj2ZpKDZ4MQ767F9UnQWpHZpawaMuqSepF1grLW32CGU%2FvRtQXlARf92Fi3HqrF2foXhumCVN7uS5nDAyMLDnbyuPMmBv0p00owLhPFf%2F7b7nkt5Gw6qzn3gRp3EIy9Y34iElMwc3lvoigmSuAzUea9qtC%2F2GFnYQVUOONOjPgTgDjAIl%2BxiLRWQ7TRfIlzaAx7xgAtI%2FxFVbCsJfeivTCt3k220nvVr3OKwcZLCYcsUThTGeGgSOkClwEw0tw0ATQ23faC5EDpJSyd1RysE8S1VsSkBl4cwv1bvp9WjEF146gUhbcqDIjb6jFYHQ2LMjIENAO%2Bn9fYXRv1cc0eNtc1TymYjNahv57GGQQ7kWBl%2BZ9eHD%2BFWLNfRUShWbKXdGMomlLmwQZ9m1k4IPEuvasYB0jcGsu5UXwxkF2GNSGSnUWjWNB5rXSQKihCJOBV46f5hXU0kw%2B8LG0wY6pAEH0b%2B2z5EdZEl%2FpW9ehNq6YC2D3NtiWEnCvkCWJPxqhlG95bkpj5l5Q2QPbROwMzjuryteuiox7Sli8YnblLLkq9qBPuS1Ozo2prS5VLmX9aFXguqqyGbEgMKY4waj%2BaZsCy%2BVSwFbICvnYvWwTiwOVV96BbDH3kJR5%2BxvE%2FYXWMeexY9NhEh%2BEzZflx6tmJnyxB2djp%2FtPXgQDS5MTNRzEzKEzg%3D%3D&X-Amz-SignedHeaders=host&X-Amz-Signature=e651b1753b549903a2cb7bc08753336bf41a1c18461ff1f1abafdbf298dcec8b
Roame is a flight and hotel search engine for points and miles that helps travelers find deals in seconds. The company is looking for a senior engineer to lead its engineering team and drive its move to a modern Go service architecture.

Keyv and friends compromised in active Shai-Hulud supply chain attack

https://cdn.prod.website-files.com/642adcaf364024552e71df01/693ab1d85c603f939d94362f_platform%20(1).png
Attackers compromised the GitHub account of keyv's maintainer, injecting credential-stealing malware into keyv and other caching utilities with 2 billion monthly installs. The malware harvests secrets from the victim's environment and exfiltrates them to a public GitHub repository.

Why etymologies matter: How tracing words can illuminate history (2024)

Ten advances in mathematics and theoretical computer science

Why Large Language Models Fail at Tabular Prediction

https://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png
Large language models fail at predictive analytics over tabular data due to their inability to handle high dimensionality. In low dimensions, LLMs behave like local, distance-based methods, but in higher dimensions, their capability dissolves and no classical model can replicate their predictions.

Devtools must be open source

https://blog.exe.dev/static/og-card.png
Software engineers used to rarely write personal software, but now agents can easily manage customizing software, making it easier to personalize and maintain. Agents can build and manage custom software with minimal programming, allowing for more efficient and flexible use of software.

Amazonian civilization had estimated 3M people in 3% of forest area

FFmpeg 9.0

https://opengraph.githubassets.com/b516399b918faa2b0f722e835527f98321d3e21b010fb4f63e6e9a2baf88dd25/FFmpeg/FFmpeg
15 lines (11 loc) · 829 Bytes ┌────────────────────────────────────┐ └────────────────────────────────────┘ We hope you will like this release as much as we enjoyed working on it, and

Mosh in a Lift (2012)

CC-BY-SA] [URL: gopher://sdf.org:70/0/users/irl/blog/2012-08-22-mosh-in-a-lift.md] Every quater, me and my flatmate pay about _80 to the Aberdeen City Council for maintainence of the lifts in our building. As I write this, I am trapped in one of these lifts. It appears that none of the buttons work except for the alarm button. You'd think this would connect you to an operator somewhere ...

There Will Come Soft Rains (1950) [pdf]

Nobel Disease

Nobel Prize winners sometimes hold scientifically unsound ideas, often due to feeling empowered to speak on topics outside their expertise. This phenomenon, known as "Nobel disease," highlights the importance of critical thinking and domain-specific expertise.

Safe Lock-free Primitives with iceoryx2's ByteAtomic

https://ekxide.io/__og-image__/static/blog/byte-wise-atomic-wrapper-to-prevent-ub/og.png
In multithreaded programming, a common scenario involves multiple threads reading from and modifying shared data concurrently. If this read and write operations are not atomic, a data race occurs. In languages like Rust and C++, which have almost the same memory model, this results in undefined behavior. To prevent this, locks can be used to protect the data from being modified while it is ...

That time when I failed the Microsoft interview

https://ochagavia.nl/images/headshot.jpg
The author, a Computer Science student, applied for a Microsoft internship in 2015 after contributing to open source projects. They received a phone interview invitation, prepared using a book on Big Tech interviewing practices, but displayed honesty when unable to solve a brain teaser, which may have affected their chances.

Twenty Years of Pandoc

John MacFarlane created the document converter pandoc in 2006 using the Haskell programming language, which he chose for its strong type system and clean representation of structured documents. Over the next 20 years, pandoc evolved into a powerful tool supporting 51 input and 76 output formats, with over 600 contributors and a reputation for reliability and determinism, despite the potential ...

Homebench – Benchmark local LLMs for speed, memory, and quality

https://camo.githubusercontent.com/2f35bc00665070bd6a3189bcff8de41f11fc182c233250fbf8850cd12a2b52f6/68747470733a2f2f63646e2e6a7364656c6976722e6e65742f67682f64617669642d672d333635342f686f6d6562656e6368406d61696e2f646f63732f64656d6f2e737667
homebench is a local-first tool that benchmarks LLMs on your laptop, measuring speed, memory, and quality. It provides a live leaderboard and supports various models and providers.

Ask HN: Who is hiring? (August 2026)

Bucket Robotics: Senior Fullstack Engineer, San Francisco, CA (On-site) - Own massive surface area, build customer-facing app, scale internal tooling, and design data pipelines. Product-minded engineer with systems thinking and curiosity.

Archaeologists Find Ancient Glyphs in the Amazon

Please enable JS and disable any ad blocker

Smaller, faster, safer: running Kimi and GLM at scale

https://blog.cloudflare.com/_image?href=https%3A%2F%2Fblog.cloudflare.com%2F_emdash%2Fapi%2Fmedia%2Ffile%2F01KZ1PVA5SC9J87PS1R2B97N1F.png&w=1999&h=1125&f=webp&fit=cover&position=center
Workers AI optimizes large models on GPUs by quantizing the KV cache to FP8, compressing model weights to INT4, and protecting the cache with integrity checks. These techniques enable efficient serving of models like Kimi K-series and GLM with no change in accuracy, supporting more customers at lower costs.

Learning-Rust.Github.io: Rust Programming Language Tutorials for Everyone

https://learning-rust.github.io/og.jpg
Learning Rust - Rust Programming Language Tutorials for Everyone!

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

https://substackcdn.com/image/fetch/$s_!59Bc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98d2552b-d84c-4385-9602-bf344dca2ebe_1000x727.png
MiniMax H3 is a next-generation open-weights video model that generates 2K video with real stereo sound up to 15 seconds. It combines text, images, video, and audio inputs with a prompt to create a single video clip, reducing five tasks into one model.

Beauty in my backyard

https://worksinprogress.co/.netlify/images?url=https%3A%2F%2Fassets.worksinprogress.co%2Fwp-content%2Fuploads%2F2026%2F07%2F471416621_122181002012057398_8649519096196337300_n.jpg&w=1080&h=1080&q=65&fit=cover
The theory that people oppose development due to ugly buildings is not supported by evidence, as many beautiful developments still face opposition, and the timing of zoning restrictions and modernist architecture does not align. Instead, the rise of conservationism and restrictive zoning is likely due to the growing opposition among people who do not live near developments, who are motivated ...