Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How about Dask - which is fairly production grade and has experimental Arrow integration.

https://github.com/apache/arrow/blob/master/integration/dask...

Dask deploys pretty well on k8s - https://kubernetes.dask.org/en/latest/



Dask does not have "experimental Arrow integration". It supports using Arrow to read Parquet files but no Arrow-based computational functionality.


Thanks for the clarification, Wes!

Semi-related question: How do you expect Arrow to be integrated to the larger data science landscape?

Will it mostly be used as a go between format? Will new libraries using it internally and old libraries just reading and translating it to a native format? Do you think established libraries will change their back-end to arrow? Is that even feasible with e.g. Pandas (or are you too far from their governance now to say)?


Too big of a discussion for Hacker News! Come on dev@arrow.apache.org if you want to talk about it


thanks for correcting me. i was not aware of this nuance. would you be open to posting quick thoughts here for the rest of us ?


No. If you want to talk about it come on the Apache Arrow mailing list


Although I know and work with some of the contributors from this project, I have no real world experience with Dask or Python data science tools in general.

Thanks for the link. I will read about their Kubernetes support.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: