Virtual Dataset basics: Controlled sharing of data across Dataspaces
Author: John Carter is a Sr Principal Product Manager at Anaplan.
As enterprises expand connected planning across finance, workforce, supply chain, and beyond, the complexity of data orchestration grows rapidly. More data sources, more integrations, and more users all increase the risk of losing visibility and control.
Dataspaces in Anaplan Data Orchestrator (ADO) unlock the ability for enterprises to organize data in ADO into clear, governed boundaries. Virtual datasets allow common data to be shared across those dataspaces with control. In this article we will cover the basics of Virtual Datasets, how they are created, how they can be used and how to remove sharing.
What are Virtual Datasets?
A Virtual Dataset provides read access to the content of a dataset from its home dataspace into any neighboring dataspace in the same region.
Figure 1: Virtual Dataset terminology
The shared dataset may be a source dataset or the result of a sequence of transformation views. It is not possible to share a dataset that is a virtual dataset or is based on one. This ensures the original data remains under the control of the integration administrators of the home dataspace as well as preventing circular references.
Where do Virtual Datasets appear in ADO?
Inventory pages for the different types of datasets – source, transformation view and virtual — are grouped under the new main menu item — All Datasets.
Figure 2: All Datasets menu
Virtual Datasets appear within their own inventory page and are classified as a different type of dataset in the neighboring dataspaces in which they appear. Their default name is based on the original dataset name appended with the name of the home dataspace. They can be renamed.
Figure 3: Virtual Dataset inventory page
The overview page shows details of how many virtual datasets there are in a dataspace and how many source dataset and transformation view based datasets have been shared with other dataspaces:
Figure 4: Virtual Datasets in the overview
Selecting the “shared” and “virtual dataset” link opens the relevant inventory page with the appropriate shared filter selected:
Figure 5: Shared filter for Datasets
Virtual Datasets also appear on the map and fall into the same swim lane as source datasets. Color coding allows the user to identify if they are a shared source dataset or a shared transformation view. They also have their own icon.
Note: Our recommendation is that transformation views are shared rather than sharing source datasets directly, even if the transformation view is a direct copy of the original source dataset. This reduces future maintenance if the data extract related to the source dataset needs to be updated in the future.
Figure 6: Virtual Datasets on the map page
How are Virtual Datasets created?
Datasets are shared from their home dataspace keeping control in the hands of Integration Administrators for that dataspace. The Integration Administrator must have access to both the home dataspace in which the dataset to be shared is based and the dataspace to which it needs to be shared where is will appear as virtual.
You can share a selected list of datasets with a neighboring dataspace using the sharing icon on the inventory page toolbar. A single dataset can be shared with multiple neighboring dataspaces from the right-hand panel.
Figure 7: Sharing a Dataset to create a Virtual Dataset
The right-hand panel shows the list of dataspaces to which the dataset has been shared.
A dataset being shared is a property of the dataset and not related to the Integration Administrator that created the share, so sharing does not stop if the Integration Administrator has access to a dataspace removed.
How can Virtual Datasets be used?
Virtual Datasets can be used like a transformation view – as the source of another transformation view, model link or pipeline. They cannot be used as a pipeline sink nor their contents updated or amended.
What happens when a Virtual Dataset is removed
A dataset can be unshared from within its home dataspace:
Figure 8: Remove sharing from within the home Dataspace
Or access to it can be removed from within the neighboring dataspace to which it has been shared:
Figure 9: Remove sharing from within the neighboring Dataspace
In both cases any dependent objects (transformation views, model links or pipelines) will become invalid.
Conclusion
Virtual Datasets allow both source and derived data to be surfaced to neighboring dataspaces for use in transformation views, model links and pipelines. The Integration Administrators of the home dataspace retain control of that data.
In the next article we will look at various use cases that virtual datasets allow us to support in ADO:
- Sharing common data
- Masking sensitive data
- Sharding central data to separate entities
- Ensuring separation of duties for external data
Learn more
To explore Virtual Datasets and Dataspaces and Anaplan Data Orchestrator in more detail:
Questions? Leave a comment!
Virtual Datasets are available for early access with General Availability planned for the end of July.