Dealer: End-to-End Data Marketplace with Model-based Pricing

by   Jinfei Liu, et al.

Data-driven machine learning (ML) has witnessed great successes across a variety of application domains. Since ML model training are crucially relied on a large amount of data, there is a growing demand for high quality data to be collected for ML model training. However, from data owners' perspective, it is risky for them to contribute their data. To incentivize data contribution, it would be ideal that their data would be used under their preset restrictions and they get paid for their data contribution. In this paper, we take a formal data market perspective and propose the first enD-to-end data marketplace with model-based pricing (Dealer) towards answering the question: How can the broker assign value to data owners based on their contribution to the models to incentivize more data contribution, and determine pricing for a series of models for various model buyers to maximize the revenue with arbitrage-free guarantee. For the former, we introduce a Shapley value-based mechanism to quantify each data owner's value towards all the models trained out of the contributed data. For the latter, we design a pricing mechanism based on models' privacy parameters to maximize the revenue. More importantly, we study how the data owners' data usage restrictions affect market design, which is a striking difference of our approach with the existing methods. Furthermore, we show a concrete realization DP-Dealer which provably satisfies the desired formal properties. Extensive experiments show that DP-Dealer is efficient and effective.


Model-based Pricing for Machine Learning in a Data Marketplace

Data analytics using machine learning (ML) has become ubiquitous in scie...

On the Fairness of Quality-based Data Markets

For data pricing, data quality is a factor that must be considered. To k...

Privacy Accounting and Quality Control in the Sage Differentially Private ML Platform

Companies increasingly expose machine learning (ML) models trained over ...

Lifelong DP: Consistently Bounded Differential Privacy in Lifelong Machine Learning

In this paper, we show that the process of continually learning new task...

Privacy Budget Scheduling

Machine learning (ML) models trained on personal data have been shown to...

A Survey of Data Pricing for Data Marketplaces

A data marketplace is an online venue that brings data owners, data brok...

Many Episode Learning in a Modular Embodied Agent via End-to-End Interaction

In this work we give a case study of an embodied machine-learning (ML) p...

Please sign up or login with your details

Forgot password? Click here to reset