Monday, August 18, 2014

When is the BEST time for a Data Quality Review? | Roshan Joseph (via LinkedIn)

Follow the LinkedIn discussion

My comment

While position 5 (NOW!) is the "correct" answer, I like to add "Merger & Acquisition" as a triggering event (variation / combination of pos. 1 to 4).

With an upcoming M&A transaction, a data quality review prepares for the audit that is an indispensable part of the due diligence. Both (all) involved organizations should undergo a data quality review to especially know about the mergability of the parties' data before taking the final decision.

Monday, June 2, 2014

Tool to Track Which Databases Keep Customer Data | LinkedIn Group: Master Data Management Pros

Follow the LinkedIn discussion

My comment

I suggest you to use a professional data and process modeling tool suite.

The data modeling tool component will allow you to have an inventory of the data, i.e. which fields (particularly: customer data) reside in which database. Typical use of the data modeling tool could be:
  • Reverse engineer each database, i.e. automatic transfer of the database structure to a graphical/textual representation in the data modeling tool.
  • (Since even (semantically) same fields will have different physical names in different databases...) Link synonyms to a common business name, e.g. "cust_name" and "cli_nam" could both represent "customer name". 
  • Add/modify any other crucial description that may be missing/incorrect.
  • Integrate the database models into subject areas (A subject area will give you the synchronized business view of how e.g. a customer is - and perspectively should be - described in your organization, e.g. by customer first-name, customer family-name, customer date-of-birth, etc.)
The process modeling tool component will provide you a graphical/textual representation how the database fields "flow" through your organization, i.e. which fields are included in the "input" and/or "output" data flow(s) of a process (program, module, dialog,...).

Ideally, the tool suite will be integrated, i.e. database fields that are captured in the database reverse engineering step (using the data modeling tool) can be linked to the fields found in the analysis of the data flows (using the process modeling tool) and vice versa.

How you apply the modeling tool suite in detail will certainly depend on the mid and long-term goals of your organization, e.g. merging/replacing application systems, evaluating new software packages, changing platforms, going mobile etc.

Considering any of these targets combined with your initial question, I recommend you to check out the SILVERRUN Professional & Enterprise Series at www.silverrun.com . (In the spirit of full disclosure: I represent Grandite, the maker of the SILVERRUN tools.)

Please do not hesitate to contact me for further information directly, you will find my coordinates in "Contact Info" of my LinkedIn profile.

Sunday, May 25, 2014

How Does the Database Influence the Data Modeling Approach? | LinkedIn Group: Data Modeling

Follow the LinkedIn discussion

My comment

Based on the 4 primary steps suggested by Rémy [Fannader] (even considering that each of us may have slightly different convictions how to exactly mark off these steps against each other), the answer should be:
  • Conceptual: not influenced by the target database 
  • Normalized logical: not influenced by the target database
  • Denormalized logical: only influenced by the architecture / type of concepts that the database supports (doesn't support), e.g. nested table, materialized view
  • Physical: completely influenced by the target database, e.g. physical names, database-specific storage parameters.
On an additional note: To follow the above 4-level procedure effectively and efficiently, it is indispensable to use a professional data modeling tool that not only, but particularly allows to
  • Define, keep and maintain the above development levels
  • Propagate (cascade) applicable modifications to the next level(s)
  • Generate the DDL from the physical level. 

Wednesday, May 7, 2014

How to Identify Parent / Child Role of an Entity in a Data Model Diagram Using "Information Engineering" Notation | LinkedIn Group: Data Modeling

Follow the LinkedIn discussion

My comment

Presuming that many-to-many relationships have been resolved as (binary) one-to-many relationships, there are two ways to communicate / express relationships, as 
  • One-to-many relationships (or optional-one-to-one)
and/or
  • Parent-child relationships [Child entity is the side of the relationship where the foreign key (constraint) will be added.]
However, provided that the integrity of the model has been positively verified, these two ways are synchronized, i.e. the role of the parent entity and of the child entity in a given one-to-many (or optional-one-to-mandatory-one) relationship can be derived following the rule "a mother can have many children, but a child has a maximum of one mother". 

The exception to this rule is an optional-one-to-optional-one relationship which needs further specification about the parent or child role of the participating entities.

For this latter case (or an interim state where the integrity of the model has not been verified yet), an Entity-Relationship diagram in Information Engineering notation only expresses the multiplicities of a relationship, but does not offer any "standard" indication about the parent / child role of an entity in a relationship.


Therefore, I suggest to use a data modeling tool that allows you to e.g. additionally display the name of the child direction close to the respective side of the relationship connector.


Once the model is verified and ready to generate foreign keys, the latter ones will graphically identify the child role of an entity in a relationship.