Python Bytes Podcast Por Michael Kennedy and Calvin Hendryx-Parker capa

Python Bytes

Python Bytes

De: Michael Kennedy and Calvin Hendryx-Parker
Ouça grátis

Ofertas da temporada | R$ 0,99/mês por 3 meses

R$ 19,90/mês após 3 meses. Confira termos e condições
Python Bytes is a weekly podcast hosted by Michael Kennedy and Calvin Hendryx-Parker. The show is a short discussion on the headlines and noteworthy news in the Python, developer, and data science space.Copyright 2016-2026 Política e Governo
Episódios
  • #498 A Tiny Episode
    Sep 29 2026
    Topics covered in this episode: MemTensor / MemoryOS PyPI package hijacked via a malicious build backendTinyMongoJev: what to knowOne innocent dict read makes attribute access permanently slowerExtrasJokeWatch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk PythonConsulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedInCalvin: Mastodon / BlueSky / X / LinkedInShow: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: MemTensor / MemoryOS PyPI package hijacked via a malicious build backend On Sept 23 an attacker published backdoored MemoryOS 2.0.34 on PyPI and three bad versions (0.1.21, 0.1.23, 0.1.25) of MemTensor's OpenClaw plugin on npm. PyPI had no clean release that day, so 2.0.34 was the newest.They pushed commits to MemTensor's own GitHub Actions release pipelines. On PyPI that was a custom Poetry build backend, and on npm a tweaked validation script. Both used BASH_ENV to hand the publish token to the attacker before the real publish ran. SafeDep couldn't confirm how the attacker got push access.Runs on import, not install: A Go implant called sckit starts when the library loads, so --ignore-scripts won't save you.It harvests credentials from your home directory (npm and PyPI tokens, GitHub tokens, SSH keys, cloud CLI tokens, .env files) and sends them to skyleen[.]fr servers.It's a worm: It uses stolen tokens to copy itself into other repos and packages, so the victim list could grow.If you installed it: Downgrade to MemoryOS 2.0.33 (plugin 0.1.20) and rotate every credential reachable from $HOME. Also kill any running sckit stage0 process and check repos you can push to for a stray runtime-update.yml workflow or .sckit/ directory. Michael #2: TinyMongo Want to use a MongoDB data interface, but swap out the storage engine? Memory for testing/cachingJSON/TinyDB simple JSON filesSQLite for durable, high-perf reads with WALSQLIte shared for high write appsDuckDB + Parquet for analytics appsPostgres + MariaDB for multi-machine client/serverGreat for teaching, examples, and simple deploymentsAmazing story of paired AI development Will completely run talkpython.fm after weeks of shared work together (in SQLite mode). Calvin #3: Jev: what to know What it is: Jev is a model from TypeSafe AI that answers with typed results (yes/no probabilities, scores, picks from your options) instead of prose. Real Python published a hands-on tutorial on 2026-09-24 and the buzz on hacker news is almost deafening.It's proprietary: Jev is a hosted, closed-weight model. There are no weights to download and no self-hosting. Everything called "open Jev" is an independent reimplementation, not TypeSafe's model.Your data leaves your machine: Every call sends your input text to a third-party API. In the tutorial that path goes through OpenRouter to TypeSafe. Think twice before sending customer messages, tickets or anything sensitive.Cost and stability are open questions: The tutorial calls Jev "cheap, but not free" and says it's fast and cheap "at the moment." It also says whether that stays true is "something to keep an eye on."Credit to Real Python: It's a good, practical intro. It shows the Noul, Score and Choice primitives, and its point that instruction wording matters more than thresholds is useful advice for any model. The tutorial itself says similar results are possible with a well-prompted LLM.Open options to look at instead: JevK5 (https://github.com/allebee/jevk5): Apache-2.0 weights and code, 4B or 9B parameters, and it accepts TypeSafe-style requests.SemIf, formerly OpenJev (https://github.com/TheoLeeCJ/openjev): MIT-licensed, small models, and it can run CPU-only.openjev-sglang (https://github.com/ekzhang/openjev-sglang): a Jev-compatible endpoint running Qwen3.6-35B-A3B, but no license is stated, so check before commercial use.The catch: These copy Jev's interface, not its model or training. Results will differ, and I haven't run any of them. Benchmarks are self-reported, and JevK5 is English-only. Michael #4: One innocent dict read makes attribute access permanently slower Timofei Ivankov benchmarks a CPython internals surprise: since 3.11, attribute access skips the instance dict entirely. A specialized opcode reads the attri.bute at a fixed byte offset in the object's inline values array. Read obj.__dict__ once, though, and the dict gets materialized, the object loses that specialized path for the rest of its life, and a million-iteration loop goes from 33 ms to 51 ms on CPython 3.14. vars() and copy.copy() trigger the same thing, so a debugging print or a shallow copy in code touching your hot objects quietly makes every later ...
    Exibir mais Exibir menos
    33 minutos
  • #497 Faster than light profiling
    Sep 23 2026
    Topics covered in this episode: Tachyon: A sampling profiler ships in Python 3.15's stdlibPython Workers are now generally available on CloudflareFlet 1.0 - build cross-platform apps in Pythonmarimo-book: Build static books from marimo notebooksExtrasJokeWatch on YouTube Sponsored by Logfire from Pydantic: pythonbytes.fm/logfire Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedInCalvin: Mastodon / BlueSky / X / LinkedInShow: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Michael #1: Tachyon: A sampling profiler ships in Python 3.15's stdlib Python 3.15 adds the profiling package per PEP 799: profiling.tracing (where cProfile moved) and profiling.sampling, the new sampler called Tachyonpy-spy and Austin exist but copy raw interpreter bytes with no API, so every CPython release risks breaking them; one in the stdlib is a contract to stop breaking profilersDefaults: 1 kHz, main thread, wall clock, and a -live top-like view for poking at a slow serverOutput is flexible: pstats, -flamegraph, -diff-flamegraph against a baseline, -heatmap on source lines, -opcodes for specialized bytecode, -gecko for Firefox Profiler with GIL and GC markersProfiling modes: wall, cpu, gil (which function is starving my other threads?), and exception, plus -async-aware to see the task graph instead of just select(), -all-threads, and -subprocesses forking a profiler per childNear-zero overhead for production; guidance is 10-30 second windows on representative load, and free-threaded builds divide the rate by thread countAttach to a running PID, same minor version only; ptrace permissions are the main friction. A 3.14 backport already exists on GitHubCaveat: it only sees Python frames, so 90% in calculate() hides NumPy underneath. For native stacks there's Cronon from HRT, 200k samples/sec over DWARF, not yet open source Calvin #2: Python Workers are now generally available on Cloudflare Python Workers are out of beta - now GA, "first-class" language on Cloudflare's Developer PlatformNo more manual JS interop: bindings (queues, R2, D1, Durable Objects) now work natively in Python, e.g. self.env.QUEUE.send({...})Runs on Pyodide (WASM-compiled Python), with real TCP socket support for DB connectivityFrameworks supported: FastAPI, Django, Flask; AI libs like OpenAI SDK, LangChain, MCPUnderlying platform work formalized as PEP 783 (PyEmscripten), after a year of discussionBottom line: write real Python on Cloudflare's edge, no JS glue code required Calvin #3: Flet 1.0 - build cross-platform apps in Python Flet hits 1.0 - build Flutter-backed apps from pure Python, no frontend experience neededOne codebase targets six platforms: iOS, Android, Windows, macOS, Linux, web150+ built-in UI controls, plus support for custom controls / wrapping Flutter packagesMobile now supports real Python packages: NumPy, pandas, Pillow, cryptographyComes with pytest-based UI testing and an MCP integration for AI coding assistantsMilestone lands 4+ years after its first PyPI release (Sept 2022) - signals "production ready," not experimental Michael #4: marimo-book: Build static books from marimo notebooks marimo-book is a Jupyter-Book-style static site generator built specifically for marimo .py notebooks. It ships polished multi-page sites with Material for MkDocs theming, full-text search, dark mode, and code copy, plus a content-hashed incremental build cache that drops rebuilds from 100+ seconds to roughly 3 seconds on real books. Standout extras include anywidget rendering without a kernel, static reactivity for discrete sliders via pre-rendered lookup tables, an opt-in WASM/Pyodide mode per chapter, and per-chapter launch buttons. If you've wanted to publish a marimo notebook as a real book or course site without hosting a kernel, marimo-book gives you the static, searchable, fast-loading output you'd expect from Jupyter Book.Alpha (0.1.x), but in production: pin marimo-book>=0.1.5,<0.2; the book.yml schema is stable for v0.1, and dartbrains.org is a real-world user.Two-stage build by design: a marimo-aware preprocessor emits plain Markdown + inline HTML, then mkdocs (Material today, zensical tomorrow) renders it. Not a mkdocs plugin, so the shell stays swappable.Interactive widgets without a kernel: anywidget Canvas/Three.js/Plotly mounts render statically, and mo.ui.slider with explicit steps gets pre-computed as a static lookup table.WASM escape hatch per chapter: set mode: wasm and the chapter routes through marimo's MarimoIslandGenerator, shipping the marimo runtime + Pyodide bundle for full reactivity where you need it.Per-chapter launch buttons and extras: readers can jump to molab, GitHub, or a downloaded .py; optional [social], [...
    Exibir mais Exibir menos
    27 minutos
  • #496 A lake house in Seattle
    Sep 15 2026
    Topics covered in this episode: Pandas Should Go ExtinctPydantic-pint puts real-world units in your Pydantic modelsHow Libraries Run Rust Inside Python (With PyO3)AWS acquires DuckLabsExtrasJokeWatch on YouTube Sponsored by Logfire from Pydantic: pythonbytes.fm/logfire Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedInCalvin: Mastodon / BlueSky / X / LinkedInShow: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Pandas Should Go Extinct Pandas' slowness pushes teams toward "Big Data" tools (Spark, Databricks) they don't actually need — most workloads never hit true Big Data scaleAmazon Redshift telemetry: ~95% of tables are under 100GB, ~87% of queries touch 80GB or less — that's "Medium Data," not Big DataPolars and DuckDB fill that gap: single-machine, fast, no cluster required1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memoryOn a real-world NYC taxi dataset (3GB parquet), pure DuckDB ran 2x faster than pure Pandas while using a fraction of the RAMBonus: Apache Arrow lets you pass data between Pandas/Polars/DuckDB with zero copying, so trying them out doesn't mean a full rewrite Michael #2: Pydantic-pint puts real-world units in your Pydantic models Pydantic-pint bridges Pydantic and Pint so models can validate physical quantities like 4m or 12 meters instead of bare floats. Fields annotated with PydanticPintQuantity parse user input, convert between compatible units, and serialize quantities back out as strings. That closes a real gap for anything consuming API payloads, config files, or sensor data with measurements, letting you enforce units at the validation boundary instead of hoping every caller remembered them. via PyCoder's Weekly newsletterUnit mix-ups have literally crashed spacecraft; now your Pydantic models can refuse them at the door.Annotate a field as Annotated[Quantity, PydanticPintQuantity('km')] and inputs like 12 meters arrive auto-converted to kilometersValidation covers string, numeric, and quantity inputs, and model_dump_json serializes quantities as readable unit stringsInstallable from PyPI as pydantic-pint, MIT licensed, with docs at pydantic-pint.readthedocs.ioEarly-stage solo project at version 0.4, so API stability and maintenance are open questions worth discussing Calvin #3: How Libraries Run Rust Inside Python (With PyO3) Pydantic v2's validation core (pydantic-core) is Rust under the hood, built with PyO3 — this post shows how that bridge actually works via a small hand-built JSON parserFour steps to get Rust into Python: write a normal Rust module, annotate with PyO3 macros (#[pyfunction], #[pymodule]), compile/install with maturin, then just import itThe parser builds a Rust tree first — Python never touches it until the boundary crossingKey insight: converting the Rust result into Python objects (.into_pyobject) is often the expensive part, not the parsing — 100,000 JSON values means ~100,000 Python objects built after parsing's already doneErrors cross the boundary too: Rust's typed errors convert into real Python exceptions (ValueError, FileNotFoundError) via From/?, so callers get clean Python semanticsTakeaway for anyone porting Rust in: if you're returning a scalar, don't sweat it; if you're returning a big structure, profile the boundary — that's the real cost, not the algorithm Michael #4: AWS acquires DuckLabs Thank you Dylan McConnell. What does this mean for the DuckDB ecosystem? DuckDB is the open-source in-process analytical SQL engine. MIT licensed. The IP is not owned by any company - it's held by the nonprofit DuckDB Foundation, which was created when the team spun out of CWI Amsterdam. Peter Boncz, the CWI representative on the Foundation board, describes it as the entity that holds all IP of open-source DuckDB. DuckLabs (ducklabs.com) is the company, formerly branded DuckDB Labs. Founded a little over five years ago by Hannes Mühleisen and Mark Raasveldt to give the DuckDB team a stable long-term home, bootstrapped deliberately instead of taking VC, grown to 30+ people in Amsterdam, funded by support and feature-prioritization contracts. It employs the core devs. It does not own DuckDB. DuckLake is one of three projects DuckLabs builds, what they call the Duck Stack: DuckDB, DuckLake, and Quack. DuckLake is the lakehouse format that puts catalog metadata in a SQL database instead of in files on object storage. Quack is newer - an RPC-style protocol that turns DuckDB into a client-server system where both ends are DuckDB instances, slated to stabilize in DuckDB v2.0 in September 2026. MotherDuck is a separate Seattle company, Jordan Tigani's, selling serverless ...
    Exibir mais Exibir menos
    33 minutos
adbl_web_anon_alc_button_suppression_t1
Ainda não há avaliações