Author: John Carter is a Sr Principal Product Manager at Anaplan.
Organize, share and control access to sensitive data
In our previous article here we focused on the basics of Virtual Datasets, what they are, where they are found in the Anaplan Data Orchestrator (ADO) UI, how a Virtual Dataset is defined, how it can then be used and what happens when a Virtual Dataset or its source is removed.
In this article, we will review four examples of how Virtual Datasets can be put to work as part of an Anaplan solution.
- Sharing common data
- Segregating sensitive data
- Sharing central data to separate entities
- Ensuring separation of duties for external data
Sharing common data
Virtual Datasets allow data from a dataset in its home dataspace to be surfaced in other neighboring dataspaces within the same region. The dataset that is shared can be a source dataset or one derived using a transformation view. The shared dataset is known as a virtual dataset in these neighboring dataspaces and can be used like other datasets within that dataspace. A virtual dataset or one derived from it with transformation views cannot be shared. This prevents circular references and ensures control of the data remains with the Integration Administrators of the home dataspace of the shared dataset.
This allows data that is used across many applications to be shared from a central point, removing duplication, ensuring consistency and one source of the truth.
Figure 1: Share common data
In this example, the Cost Centre details extracted from the external source system into the Ingest Dataspace are shared with dataspaces for several solutions: FP&A, Product Profitability and Long-Range Planning.
Note that the source dataset itself has not been shared, rather a transformation view based on that source dataset. Adding this transformation view, even if this is initially a trivial select all columns and rows affair, provides flexibility for future changes. For example, if the external source of the data changes or the structure and content needs to be restructured in some way. Such changes are then accommodated by remapping the shared transformation view and updating its configuration if necessary.
Segregating sensitive data
It is typical to store sensitive data in separate dataspaces in ADO to allow access to that data to be kept to a minimum. For example, the data needed for Workforce Planning will include data at a level that allows personal information to be identified. It is likely that such applications will make use of common shared data, such as the Cost Centre example illustrated above. It is also possible that some of the data in this restricted dataspace is needed for other solutions either in an anonymized form or at an aggregated level.
Sticking with the example of Workforce planning, although the information by employee is not required for other applications, figures such as the total headcount and total workforce costs by cost center are. Virtual datasets provide an ability to achieve this.
Figure 2: Segregating sensitive data
In Figure 2, data from the external system is extracted into the dataspace dedicated to Workforce planning. These details are not shared directly; they are filtered and aggregated using one or more transformation views to remove and personal information and to ensure that the data cannot be associated with any individual. This derived dataset is shared with the relevant dataspaces focused on solutions that require total workforce costs and headcount.
In addition to ensuring suitable access control for sensitive data, this approach avoids the complexity and overhead of multiple extracts from the source and ensures that the data, albeit at different levels of detail, is consistent across the solutions, which is an important consideration for connected planning.
Sharing centrally held data across separate entitles
In scenarios where there is a single source application, such as the corporate ERP, and where business entitles plan discreetly, virtual datasets can also provide a suitable approach. They allow the data to be extracted from the source and suitably filtered into relevant subsets for each business entity. Each filtered dataset can be shared as a virtual dataset to the dataspace reserved for that business entity.
Figure 3: Segmented data
In Figure 3 we see an example of data extracted from the corporate ERP that needs to be split by operating company. Transformation views are used to filter the data into separate shards for Co1, Co2 and Co3. Each of these is then shared to the relevant consuming dataspace. Each operating company gets to see their own data and only their own data. Extracting the data from the corporate source into a single place ensures consistency, central control, and efficiency, while the use of Virtual Datasets ensures that the operating companies get access to only what they need.
Ensuring separation of duties for external data
The example above leads neatly into the final example use case for virtual datasets covered in this article – central control of access to an external data source and connections to it.
Within ADO it is possible to restrict access to dataspaces by user. The tenant administrator controls which users with the integration administration role have access to which dataspaces. However, once access to a dataspace is granted, a user has full access to all objects in that dataspace. That includes any defined connections. Additional data extracts can be created for that connection by any Integration Administrator with access to the dataspace.
In situations where access to the connection needs to be limited and tightly controlled, but where data from a connection needs to be accessible, virtual datasets provide a powerful solution.
Figure 4: Separation of duties
The figure above illustrates how access to connections to sources (the EDW in this example) can be limited while access to the data from them can be granted to the Integration Administrator users that require it.
The tenant administrator limits access to the ingest dataspace to the administrators for that source. In turn, those users share the required data to the appropriate dataspaces for the users that need it as virtual datasets. In this way, there is governance and separation of duties. The data source administrators remain in control of the connection and data from the EDW while the users focused on applications have all the access needed to that data.
Conclusion
Virtual Datasets allow both source and derived data to be surfaced to neighboring dataspaces for use in transformation views, model links and pipelines. The Integration Administrators of the home dataspace retain control of that data.
In this article we looked at four use cases that virtual datasets allow us to support in ADO:
- Sharing common data: ensures consistency and one source of the truth
- Segregating sensitive data: provides access control for confidential data while avoiding the complexity of multiple extracts ensuring data constancy
- Sharing central data to separate entities: allows a single extract to be partitioned across multiple entities for efficiency and uniformity
- Ensuring separation of duties for external data: allows access to the data modelers require while limiting access to the data administrators.
These demonstrate the power and flexibility offered by Virtual Datasets allowing the Integration Administrator to minimize duplication, ensure consistency, control access to and control of sensitive data, support sharded solutions and ensure separation of duties.
Virtual Datasets are in General Availability.
Learn more
To explore Virtual Datasets and Dataspaces and Anaplan Data Orchestrator in more detail: