Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Love the idea! I really see the value in shifting the conversation towards the vendor themselves being responsible for pushing the data to customers and it makes a lot of sense to do it directly from DB -> DB.

However, building a data product myself (Shipyard), we really try to encourage the idea of "connecting every data touchpoint together" so you can get an end-to-end view of how data is used and prevent downstream issues from ever occurring. This raised a few questions:

1. If the vendor owns the process of when the data gets delivered, how would a data team be able to have their pipelines react to the completion or failure of that specific vendor's delivery? Or does the ingestion process just become more of a black box?

While relying on a 3rd party ingestion platform or running ingestion scripts on your own orchestration platform isn't ideal, it at least centralizes the observability of ongoing ingestion processes into a single location.

2. From a business perspective, do you see a tool like Prequel encouraging businesses to restrict their data exports behind their own paywall rather than making the data accessible via external APIs?

--

Would love to connect and chat more if you're interested! Contact is in bio.



1. I think in practice, people are already using a mix of sources for ingestion today. It’s rare that a data team would rely on a single tool to ingest all their data – instead, they might get some data via one or more ETL tools, some data via a script they wrote themselves, and some other data from their own db. So in that regard, I don’t think a world where the vendor provides the data pipeline makes this a lot more complex.

One way we’re hoping to pre-empt some of this is by helping vendors to surface more observability primitives in the schema that they write data to. To give an example: Prequel writes a _transfer_status table in each destination with some metadata about the last time data was updated. The goal there is to decouple the means of moving data with the observability piece.

We can also help vendors expose hooks that people’s data pipelines & observability tools plug into (think webhooks and the like).

2. We don’t really – anecdotally, companies that offer data warehouse syncs tend to be pretty focused on providing a great user-experience. At the end of the day, that decision is pretty much entirely with the business. We see a pretty wide range today: some teams choose to make exports widely available, some choose to reserve it for their pro or enterprise tier, and some choose to sell it as a standalone SKU. It’s pretty similar to what already happens with APIs designed for data exports.

Would love to continue the convo and hear more about your take on obs! Sending you a note now.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: