Datalake
Tuned Global DataLake gives customers controlled access to operational and usage data generated through their Tuned Global services.
It allows customer data teams to combine music-service activity with information held across their own business, including customer, subscription, billing, CRM, telecommunications, marketing and product data.
This can provide a broader view of how users engage with both the music service and the customer’s wider product ecosystem.
For example, a telecommunications provider may use a common customer identifier to combine:
- music listening behaviour;
- subscription and package information;
- mobile-service usage;
- customer tenure;
- campaign engagement;
- support interactions; and
- other customer information held within its own systems.
This can help the provider understand the relationship between music engagement and broader customer behaviour, including retention, package adoption, service usage and product satisfaction.

DataLake is intended for organisations that require more than predefined reports or dashboards. It provides data in a form that can be queried, modelled and integrated into an existing analytics environment.
Common uses include:
- business intelligence and executive reporting;
- audience and behavioural analysis;
- customer segmentation;
- churn and retention modelling;
- subscription analysis;
- product and feature optimisation;
- content-performance analysis;
- campaign measurement;
- recommendation and personalisation workflows;
- rights, usage and operational reconciliation; and
- approved machine-learning and data-science use cases.
This document describes the standard capabilities, delivery model and principal data domains of Tuned Global DataLake.
It is intended to help technical, product, data, security and procurement teams evaluate whether the service is suitable for their requirements.
It is not the definitive schema or implementation specification for an individual customer.
Following activation, Tuned Global provides a customer-specific implementation document containing the exact datasets, schemas, field definitions, delivery configuration, identifiers, access arrangements and other technical information required to implement and operate the service.
What DataLake Provides
Tuned Global DataLake provides structured analytical data derived from the customer’s enabled Tuned Global services.
Depending on the agreed configuration, customers may receive:
- curated, analysis-oriented datasets;
- source-level datasets;
- contextual content information;
- user and subscription information;
- device and application information;
- playlist, station and tag information; and
- calculated or aggregated activity measures.
The precise datasets available depend on:
- the Tuned Global services used by the customer;
- the customer’s application and authentication design;
- the data generated by the applicable product;
- the information available from connected devices and applications;
- privacy, security and regulatory requirements;
- the customer’s rights and permissions;
- the agreed DataLake configuration; and
- any customer-specific integration requirements.
A field or dataset described in this public reference should not be interpreted as confirmation that it will be available or populated for every customer, user or event.
How DataLake Works
Tuned Global processes activity and operational records generated across the customer’s enabled services.
Relevant data is validated, transformed and made available through a customer-specific DataLake delivery.
The standard delivery model uses files in Apache Parquet format stored in an Amazon S3 location configured for the customer.
Parquet is a column-oriented data format designed for analytical workloads. It is widely supported by cloud data platforms, data warehouses, query engines and data-processing tools.
Customers can use the delivered data in several ways, including:
- querying the files directly;
- incorporating the files into an existing data lake;
- loading or referencing the data from a data warehouse;
- transforming the datasets into customer-specific analytical models;
- combining the data with other customer-controlled datasets; and
- exposing selected outputs to reporting, personalisation or data-science systems.
Data is normally delivered incrementally.
Each scheduled delivery contains newly available or changed records for the applicable period, rather than requiring the entire historical dataset to be reissued with every update.
The standard refresh cycle is daily. The applicable timing, processing window and delivery configuration are confirmed during implementation.
DataLake is downstream of Tuned Global’s operational systems. Events are first received and processed through the relevant Tuned Global services before being made available through the DataLake pipeline.
DataLake should therefore be treated as an analytical data service rather than a real-time transactional interface.
It should not be used as the authoritative source for immediate entitlement, authentication or playback decisions.

Delivery Format and Compatibility
Parquet format
DataLake files are ordinarily delivered in Apache Parquet format.
Parquet provides several advantages for analytical use:
- efficient column-based queries;
- reduced data scanning compared with many row-based formats;
- compression and storage efficiency;
- support for typed schemas;
- compatibility with distributed data-processing platforms; and
- broad support across modern cloud analytics environments.
Amazon S3 delivery
Data is delivered to an Amazon S3 location configured as part of the customer implementation.
The final access model may vary depending on the agreed architecture and security requirements.
The customer-specific implementation document confirms matters such as:
- AWS region;bucket and path configuration;
- access roles and policies;
- encryption arrangements;
- folder and partition structure;
- file-naming conventions;
- delivery frequency;
- retention arrangements; and
- operational contacts.
Analytics compatibility
Parquet files can be used with a range of compatible technologies.
Typical examples may include:
- Amazon Athena;
- Amazon Redshift;
- Snowflake;
- Databricks;
- Apache Spark;
- Trino or Presto-compatible query engines;
- cloud-native data lakes;
- business-intelligence platforms; and
- customer-built processing pipelines.
Compatibility depends on the customer’s selected technology, configuration and implementation approach.
Tuned Global provides the delivered datasets and applicable implementation information but does not, unless separately agreed, operate the customer’s warehouse, business-intelligence environment or downstream data models.
Dataset Options
DataLake can include curated datasets, source-level datasets or a combination of both.
Curated datasets
Curated datasets combine related information into analysis-oriented records.
For example, a listening-activity dataset may combine:
- a listening event;
- a user identifier;
- content information;
- playback-source context;
- subscription context;
- device information;
- listening duration;
- completion or skip indicators; and
- relevant classification or quality information.
This reduces the number of joins required for common reporting and analytics use cases.
Curated datasets are generally the most direct starting point for customers that want to build dashboards, behavioural models, customer views or business reports without reproducing Tuned Global’s underlying data relationships.
Source-level datasets
Source-level datasets expose selected underlying records with less transformation.
They provide greater flexibility for customers that wish to construct their own:
- relationships;
- data models;
- business rules;
- measures;
- historical views; and
- analytical outputs.
Source-level datasets may include separate records relating to:
- play and listening activity;
- users;
- user details;
- authentication;
- subscriptions;
- devices;
- playlists;
- radio or station services;
- tags;
- content context; and
- calculated popularity or engagement information.
Source-level datasets require a greater understanding of identifiers, relationships, record lifecycle and joining logic.
The applicable keys, relationships, field definitions, data types and joining guidance are provided in the customer-specific implementation document.
Principal Data Domains
The following sections describe the principal categories of information that may be available through DataLake.
Availability and population vary by customer implementation.
Listening and usage activity
Listening and usage datasets may include information such as:
- unique activity or session identifiers;
- user identifiers;
- event date and time;
- playback source;
- source identifier and source type;
- content identifier;
- content type;
- listening duration;
- total content duration;
- start, progress, completion or skip activity;
- stream, offline or other supported usage context;
- audio quality or asset information;
- application or device context;
- country or available location context;
- trial or subscription context; and
- other event information generated by the applicable service.
These datasets may support:
- active-user reporting;
- listening-time analysis;
- skip and completion analysis;
- content-performance reporting;
- feature and source analysis;
- subscription engagement analysis;
- product optimisation; and
- usage reconciliation.
User and identity information
User datasets may include:
- Tuned Global user identifiers;
- customer-supplied or external user identifiers;
- account status;
- registration date;
- available language or country information;
- profile relationships;
- trial indicators;
- available user attributes; and
- timestamps relating to record creation or updates.
The availability of personal or profile information depends on the implementation, applicable permissions and the customer’s service design.
Personal information is not automatically included merely because a potential field exists within a source system.
Subscription information
Subscription datasets may include:
- subscription or entitlement identifiers;
- package identifiers;
- package or service names;
- subscription status;
- start and end dates;
- trial status and expiry;
- external subscription references;
- renewal or processing information; and
- timestamps relating to subscription lifecycle events.
These datasets may support:
- paid-subscriber reporting;
- free-to-paid conversion analysis;
- package-performance analysis;
- trial analysis;
- churn and retention analysis; and
- engagement by subscription type.
Device and application information
Device and application datasets may include:
- device identifiers;
- device type;
- operating system;
- device name or model information;
- language;
- timezone offset;
- carrier information, where available;
- application identifier or context; and
- user-to-device relationships.
Device data can vary substantially between platforms and operating systems.
The information available depends on:
- application implementation;
- device capabilities;
- operating-system restrictions;
- customer permissions;
- SDK or API inputs; and
- the information provided by the originating device.
Content context
Content context may include information associated with activity records, such as:
- Tuned Global content identifiers;
- track or content title;
- artist identifier and name;
- album or release title;
- ISRC;
- UPC;
- genre;
- provider or delivery source;
- label information;
- content type;
- total duration; and
- available language or localisation information.
DataLake is not intended to provide a complete export of all catalogue metadata.
Unless separately agreed, catalogue information is generally provided where it is relevant to activity or other records included in the customer’s DataLake.
Playlist and radio information
Where the applicable services are enabled, DataLake may include information relating to: playlists;
- stations or radio services;
- names and descriptions;
- tags;
- language;
- source;
- ownership or user relationships;
- status;
- content tier; and
- creation or update information.
This may support analysis of:
- editorial performance;
- playlist engagement;
- radio or station engagement;
- discovery behaviour;
- tag performance; and
- user-generated or editorial content usage.
Aggregated measures
Where enabled, DataLake may include calculated or aggregated measures relating to content or artist activity.
Examples may include:
- plays;
- skips;
- likes;
- dislikes;
- shares;
- unique users;
- daily activity;
- weekly activity;
- monthly activity;
- total activity; and
- ranking or scoring measures.
The calculation method, applicable time windows and dataset availability are confirmed during implementation.
Representative Dataset Summary
Data domain | Typical information | Example uses |
Listening and usage activity | Play events, progress, completion, skips, source context, duration and content identifiers | Engagement, content analysis, reconciliation and product optimisation |
Users and identity | User identifiers, account state, external references and available profile attributes | Customer-360 analysis, cohorts, retention and segmentation |
Subscriptions | Package, status, trial, lifecycle dates and external references | Conversion, entitlement, revenue and churn analysis |
Devices and applications | Device, operating system, language, carrier and application context | Device segmentation, quality analysis and support |
Content context | Track, release, artist, ISRC, UPC, genre and provider information | Content performance and audience analysis |
Playlists and radio | Playlist, station, text, tags and related activity context | Editorial, discovery and source-performance analysis |
Aggregated measures | Plays, skips, likes, shares, users and time-based totals | Trending, ranking and catalogue-performance analysi |
This table is representative rather than exhaustive.
Exact schemas and field availability are documented for each customer deployment.
Identity and Data Joining
A key benefit of DataLake is the ability to associate Tuned Global activity with customer-controlled information.
Where a customer’s authentication or integration design supplies a stable customer identifier, that identifier may be made available alongside the relevant Tuned Global user identifier.
This can allow the customer to join DataLake records with systems such as:
- customer relationship management platforms;
- billing systems;
- subscription platforms;
- telecommunications data environments;
- customer data platforms;
- marketing systems;
- loyalty systems; and
- customer-support platforms.
The ability to create a unified customer view depends on the existence and lawful use of a suitable common identifier.
Tuned Global does not assume that identities can be joined in every implementation.
The identity approach is confirmed during solution design and may include:
- a Tuned Global user identifier;
- a customer external user identifier;
- a subscription identifier;
- a pseudonymous identifier;
- a household or profile relationship; or
- another agreed mapping method.

Data Refresh and Incremental Delivery
The standard DataLake refresh cycle is daily.
The exact processing and delivery window is agreed as part of implementation.
Data is normally provided through incremental files containing records that are newly available or have changed since the applicable prior processing period.
Depending on the dataset, customers may need to account for:
- new records;
- updated records;
- late-arriving events;
- corrected records;
- duplicate-delivery protection;
- historical backfills; and
- record lifecycle or status changes.
The customer-specific implementation document describes the expected processing method for each enabled dataset.
The customer is responsible for implementing appropriate downstream ingestion controls unless a separate managed integration service has been agreed.
These controls may include:
- file discovery;
- ingestion tracking;
- idempotent processing;
- duplicate handling;
- schema validation;
- error handling;
- reconciliation;
- monitoring; and
- alerting.
Data Characteristics and Limitations
Analytical rather than real-time
DataLake is designed for analytical use.
It is not a real-time event stream and should not be used for decisions requiring immediate transactional state.
Field population
Not every field will be populated for every record.
Field availability may depend on:
- service configuration;
- event type;
- content type;
- user status;
- device capabilities;
- application implementation;
- customer permissions;
- geography;
- source-system availability; and
- privacy or regulatory restrictions.
Device and location information
Device, carrier, IP-derived, location or application information may be incomplete or unavailable.
Customers should not assume that such information is precise, consistently available or suitable for safety-critical or highly sensitive decisions.
Catalogue scope
Content information in DataLake is generally contextual.
It relates to activity, playlists, stations or other records included in the DataLake.
It should not be treated as a complete catalogue export unless a separate catalogue delivery has been expressly agreed.
Calculated fields
Some fields may be derived from other records or calculated using defined business rules.
For example, a skip indicator may depend on the reported listening duration, event sequence or completion state.
The definitive calculation rules are documented in the customer implementation specification where applicable.
Schema evolution
Datasets may evolve as products, services and source systems change.
Tuned Global manages material changes through the applicable operational and change-management process.
Customers should design downstream ingestion so that it can appropriately handle agreed schema changes.
Customer interpretation
Tuned Global provides data and applicable field definitions but does not control all conclusions or models created by the customer.
Customers remain responsible for:
- their downstream business logic;
- analytical interpretation;
- customer scoring;
- automated decisions;
- model governance;
- compliance; and
- the appropriate use of DataLake outputs.
Security and Access
DataLake is delivered through a customer-specific access configuration.
The applicable controls are agreed during implementation and may include:
- restricted AWS Identity and Access Management roles;
- least-privilege access;
- encryption in transit;
- encryption at rest;
- customer-specific storage paths;
- access logging;
- environment separation;
- retention controls; and
- credential or role-management procedures.
Customers are responsible for security within their own environment once data has been accessed or copied into their systems.
This includes responsibility for:
- downstream access permissions;
- internal user access;
- storage;
- backups;
- processing;
- exports;
- retention;
- deletion;
- monitoring; and
- third-party access.
The exact allocation of responsibilities is documented in the applicable agreement and implementation documentation.
Privacy and Personal Data
DataLake may include information relating to users, devices or service activity.
The inclusion of personal data depends on the customer implementation, available data, lawful basis, contractual scope and agreed security controls.
Personal-data fields may be:
- included;
- excluded;
- limited;
- pseudonymised;
- tokenised;
- masked; or
- otherwise controlled.
Customers should identify the minimum data required for their intended use cases.
Tuned Global and the customer will confirm the applicable data treatment as part of implementation.
Customers remain responsible for their downstream use of the data, including:
- privacy notices;
- user rights;
- lawful use;
- access controls;
- retention;
- deletion;
- international transfers;
- data-sharing arrangements; and
- automated decision-making requirements.
DataLake should not be used to create sensitive or intrusive user profiles without an appropriate legal, ethical and contractual basis.
Typical Implementation Process
A standard DataLake activation may include the following stages.
Requirements confirmation
Tuned Global and the customer confirm:
- business use cases;
- required datasets;
- reporting objectives;
- identity requirements;
- historical-data requirements;
- refresh expectations; and
- downstream technology.
Data and privacy review
The parties confirm:
- required user information;
- personal-data treatment;
- excluded fields;
- pseudonymisation requirements;
- retention considerations; and
- applicable security or regulatory constraints.
Architecture and access design
The parties confirm:
- AWS region;
- storage and access model;
- IAM roles or policies;
- encryption;
- delivery paths;
- environment separation; and
- operational ownership.
Dataset configuration
Tuned Global confirms:
- enabled datasets;
- curated and source-level outputs;
- schema;
- identifiers;
- relationships;
- partitions;
- delivery cadence; and
- any agreed historical backfill.
Provisioning and validation
Tuned Global provisions the agreed delivery and provides sample or initial data for validation.
The customer validates:
- access;
- file ingestion;
- schema handling;
- identifier mapping;
- expected volumes;
- data quality; and
- downstream processing.
Production activation
Following validation, production delivery is activated in accordance with the agreed schedule.
Customer Implementation Documentation
Following activation, Tuned Global provides a customer-specific implementation document.
This document contains the detailed technical information required to ingest, interpret and operate the customer’s DataLake delivery.
It will ordinarily address the applicable items below.
Delivery configuration
- AWS region;
- bucket and path details;
- access configuration;
- encryption;
- environments;
- file-naming conventions;
- folder structure;
- partitions;
- delivery schedule; and
- retention arrangements.
Dataset specification
- enabled datasets;
- exact table or dataset names;
- fields;
- data types;
- nullability;
- formats;
- enumerations;
- field descriptions;
- calculated fields;
- deprecated fields; and
- dataset-specific notes.
Relationships and identifiers
- primary identifiers;
- foreign-key relationships;
- Tuned Global identifiers;
- customer external identifiers;
- session or event identifiers;
- profile or household relationships;
- content identifiers; and
- joining guidance.
Incremental processing
- delivery behaviour;
- insert and update handling;
- late-arriving records;
- correction handling;
- historical loads;
- duplicate handling;
- replay or redelivery procedures; and
- reconciliation guidance.
Data-quality guidance
- expected field population;
- known limitations;
- validation rules;
- event sequencing;
- calculated-field logic;
- timing considerations; and
- applicable service-specific behaviour.
Operational information
- implementation contacts;
- support process;
- incident escalation;
- delivery monitoring;
- access changes;
- schema-change communication; and
- service review arrangements.
Implementation documentation
The customer-specific implementation document, rather than this public reference, is the authoritative technical description of the customer’s DataLake deployment.
It contains the exact schemas, fields, identifiers, relationships, delivery paths, access controls and operating procedures applicable to that customer.
Customer Responsibilities
A successful DataLake implementation normally requires participation from the customer’s:
- data-engineering team;
- cloud or infrastructure team;
- security team;
- privacy or legal team;
- product or analytics team; and
- relevant business stakeholders.
Unless separately agreed, the customer is responsible for:
- providing accurate implementation requirements;
- identifying the lawful and permitted use cases;
- providing any required external identifiers;
- configuring its downstream environment;
- ingesting and storing delivered data;
- managing customer-side access;
- building customer-specific models;
- validating outputs;
- monitoring downstream processing;
- maintaining appropriate privacy controls; and
- ensuring that DataLake outputs are used appropriately.
Evaluating DataLake for Your Organisation
DataLake may be appropriate where an organisation:
- requires detailed access to its service data;
- wants to integrate music activity with enterprise data;
- operates an established analytics or data-warehouse environment;
- requires customer-specific reporting;
- wants to build behavioural or retention models;
- needs greater modelling flexibility than standard dashboards provide;
- wants to analyse listening, subscription, device and content behaviour together; or
- has data-engineering capability to ingest and govern analytical datasets.
Before activation, customers should consider:
- the decisions they want the data to support;
- which datasets are actually required;
- whether a stable common user identifier exists;
- whether curated or source-level data is preferable;
- the required history and refresh frequency;
- privacy and personal-data requirements;
- the customer’s downstream data platform;
- expected data volumes;
- security and access requirements; and
- internal ownership of the resulting data product.
Summary
Tuned Global DataLake provides structured access to analytical data generated through a customer’s Tuned Global services.
It enables technical and data teams to:
- analyse listening and service behaviour;
- combine music engagement with customer-controlled information;
- build customer-specific reporting and data models;
- understand subscription, device, content and editorial performance;
- support retention, segmentation and product optimisation; and
- integrate music-service data into a wider enterprise analytics environment.
Data is ordinarily delivered in Parquet format through Amazon S3 and refreshed incrementally on a daily cycle.
Customers may receive curated datasets, source-level datasets or both.
This public reference explains the standard product and delivery model. The exact technical configuration is documented in a customer-specific implementation document supplied as part of activation.
Frequently Asked Questions
What is Tuned Global DataLake?
Tuned Global DataLake provides structured access to usage and operational data generated through a customer’s enabled Tuned Global services.
What data can DataLake include?
Depending on the implementation, DataLake may include listening activity, users, subscriptions, devices, content context, playlists, radio services and aggregated engagement measures.
How is DataLake delivered?
Data is ordinarily delivered as Apache Parquet files to a customer-specific Amazon S3 location.
How often is the data refreshed?
The standard refresh cycle is daily, with incremental deliveries containing newly available or updated records.
Is DataLake a real-time service?
No. DataLake is designed for analytics and downstream processing rather than immediate transactional decisions.
Can DataLake be combined with our own business data?
Yes. Customers can combine Tuned Global data with CRM, billing, subscription, loyalty, telecommunications and other enterprise datasets within their own analytics environment.
Does Tuned Global ingest or enrich our enterprise data?
No. Tuned Global prepares and delivers data generated through its services. Any joining with customer-controlled datasets occurs within the customer’s environment.
What analytics platforms can be used?
Parquet files can be used with compatible technologies such as Amazon Athena, Amazon Redshift, Snowflake, Databricks, Apache Spark and other modern analytics platforms.
Are curated and source-level datasets available?
Yes. Depending on the agreed configuration, customers may receive analysis-ready curated datasets, selected source-level datasets, or both.
Can DataLake be used for machine learning?
Yes, subject to the customer’s rights, permissions and applicable contractual requirements. Customers can use delivered data within approved analytics and machine-learning workflows.
Does DataLake include a complete catalogue export?
Not ordinarily. Content information is generally provided where it is relevant to activity, playlists, stations or other included records.
Is historical data available?
Historical data may be available depending on the services used, data availability and the agreed implementation scope.
Is DataLake secure?
DataLake uses customer-specific access controls and may include restricted IAM access, encryption, logging and environment separation.
Does every customer receive the same schema?
No. Dataset availability and field population depend on the customer’s enabled services, application design, permissions and agreed configuration.
What documentation is supplied during implementation?
Customers receive an implementation specification covering enabled datasets, exact schemas, identifiers, relationships, delivery paths, access controls and operational procedures.
On this page
- Datalake