4  Guidelines for privacy-compliant results

In order to release the results from jobs in Presentation/Publication Mode, a data protection review is carried out at the FDZ in a two-stage procedure: First, a specially developed test script is used to automatically delete critical values. However, this script can only check simple output formats in Stata (e.g. for the commands tabulate oneway, summarize) and delete critical values. Afterwards, a manual check of all results is always carried out by the FDZ employees.

As a rule, the FDZ does not make any manual deletions in the result files, but only checks whether the requirements for the minimum numbers of observations are met or whether too small numbers of observations have been censored by the FDZ’s automatic check script. In exceptional cases, the FDZ undertakes manual deletions if it can be justified that a grouping of categories of a variable is not possible in order to avoid the output of too small numbers. Subsequent manual deletions by users are not possible.

You must ensure that your programmes do not generate output with too small numbers of observations. Examples for the design of the do-files are included in the FDZ templates. In the following, the most important guidelines of the FDZ for the data protection review of analysis results are presented. The following examples are intended to illustrate, on the one hand, which results are classified as questionable under data protection aspects and, on the other hand, to help you avoid the output of such results.

4.1 Number of observations in general

  • For all descriptive or multivariate analyses, always provide the number of observations they are based on. If the number of cases is missing, the corresponding results will not be released.
  • For data protection reasons, all results must be based on at least 20 observations. This applies to all statistical figures, cross-tabulations and also multivariate analyses. The minimum requirement of 20 observations of the unit under consideration applies to enterprise-level, establishment-level and individual-level data offered by the FDZ.
  • As part of the data protection check at the FDZ, log files or, if necessary, the entire job will not be released if values are contained that are based on fewer than 20 observations (individuals and/or establishments and/or enterprises).
  • For analyses of linked data containing enterprise, employer and/or employee data, the numbers of observations for individuals, establishments and enterprises have to be reported in descriptive tables.

4.2 Frequency counts

In frequency counts, values < 20 are usually not released. In cross tables, the number of cases of each individual cell must be >= 20, the total number of the table is not sufficient here. Therefore, if evaluations are based on a too small number of observations, the programmes for their generation must be adapted by you in such a way that insufficient case numbers are no longer displayed. For example, in descriptive statistics, characteristics can be summarised or excluded so that manual deletions are no longer necessary. Example 1 shows how the categories “100-499 SVB” (SVB = „Sozialversicherungspflichtig Beschäftigte” / Employees subject to social security) and “500-999 SVB” can be combined to prevent the output of values that are too small. Further instructions can be found in the FDZ template jd02_describe.do.

Example 1: East Germany

Table 4.1: Before
Works Council
No. employees (SVB) Yes No Total
1 1-4 SVB 43 1380 1423
2 5-9 SVB 39 547 586
3 10-99 SVB 594 1322 1916
4 100-499 SVB 573 175 748
5 500-999 SVB 142 16 158
Total 1391 3440 4831
Table 4.2: After
Works Council
No. employees (SVB) Yes No Total
1 1-4 SVB 43 1380 1423
2 5-9 SVB 39 547 586
3 10-99 SVB 594 1322 1916
4 100-999 SVB 715 175 906
Total 1391 16 4831

If it is not possible to group categories for important reasons, manual censoring of the results can be carried out by the FDZ team in exceptional cases. First, all values (number of observations, distribution statistics and regression coefficients) based on a number of observations below 20 have to be deleted. Second, to prevent the re-calculation of deleted values by using subtotals or marginal sums, additional values have to be deleted or rounded.

Example 2 explains the procedure of the FDZ for data protection checks of descriptive cross tables. First, values < 20 are replaced by “/” due to the insufficient number of observations (primary suppression). In order to prevent a re-calculation of the deleted value by using the marginal sum, further values are subjected to secondary suppression (“*”).

Additionally, the re-calculation across several tables must be prevented. Example 3 illustrates this case. In West Germany, the number of observations in the individual cells is >=20 in each case, but since the table for all of Germany is also shown in addition to East and West Germany, a re-calculation of the deleted values from East Germany (from example 2) is possible. For this reason, the corresponding cells in the table for West Germany are also censored (secondary suppression: “*”).

Example 2: East Germany

Table 4.3: Before
Works Council
No. employees (SVB) Yes No Total
1 1-4 SVB 43 1380 1423
2 5-9 SVB 39 547 586
3 10-99 SVB 594 1322 1916
4 100-499 SVB 573 175 748
5 500-999 SVB 142 16 158
Total 1391 3440 4831
Table 4.4: After
Works Council
No. employees (SVB) Yes No Total
1 1-4 SVB 43 1380 1423
2 5-9 SVB 3* 54* 586
3 10-99 SVB 594 1322 1916
4 100-499 SVB 573 175 748
5 500-999 SVB 14* / 158
Total 1391 3440 4831

Example 3: West Germany

Table 4.5: Before
Works Council
No. employees (SVB) Yes No Total
1 1-4 SVB 64 2461 2525
2 5-9 SVB 54 847 901
3 10-99 SVB 859 1985 2844
4 100-499 SVB 793 255 1048
5 500-999 SVB 198 22 220
Total 1968 5570 7538
Table 4.6: After
Works Council
No. employees (SVB) Yes No Total
1 1-4 SVB 64 2461 2525
2 5-9 SVB 5* 84* 901
3 10-99 SVB 859 1985 2844
4 100-499 SVB 793 255 1048
5 500-999 SVB 19* 2* 220
Total 1968 5570 7538

Germany Total

Table 4.7: Before
Works Council
No. employees (SVB) Yes No Total
1 1-4 SVB 107 3841 3948
2 5-9 SVB 93 1394 1487
3 10-99 SVB 1453 3307 4760
4 100-499 SVB 1366 430 1796
5 500-999 SVB 340 38 378
Total 3359 9010 12369
Table 4.8: After
Works Council
No. employees (SVB) Yes No Total
1 1-4 SVB 107 3841 3948
2 5-9 SVB 93 1394 1487
3 10-99 SVB 1453 3307 4760
4 100-499 SVB 1366 430 1796
5 500-999 SVB 340 38 378
Total 3359 9010 12369

Due to the risk of deanonymization, you need to report the number of observations for all subgroups, if the re-calculation of deleted values is possible. In the example above, you need to show the table for both West and East Germany, since the table for Germany allows the re-calculation of values for East Germany. If, for example, instead of an east-west comparison, you want to analyse only the construction industry, the output of all industries of the economy is unnecessary, since no re-calculation is possible.

4.3 Descriptive statistics

At first glance, descriptive statistics do not allow any conclusion to be drawn about the underlying case numbers. However, this does not mean that the output of descriptive statistics, such as means, cannot be problematic. Here, too, the principle applies that the statistics addressed are only classified as safe if the calculation basis comprises at least 20 observations. A special case is the output of descriptive statistics for dummies. In the case of binary-coded variables, the values of these are distributed over only two categories. Even if the total number of observations of a dummy variable is greater than 20, it may be that due to a skewed distribution only three persons or establishments or enterprises fall into one of the categories. In this case, the output is considered uncertain, even if the total number of observations does not directly indicate a low occupancy of a category. This is because this can be easily determined using the means. In order to be able to identify and check these cases, it is necessary to document whether a dummy variable is involved when means are displayed.

Example 4 illustrates the problem with the output of mean values. In the case of the dummy variable r61, all values had to be censored because only 12 establishments (140*0.0857143) fall into one of the two categories due to the skewed distribution.

Example 4

Table 4.9: Before
Variable Obs Mean Std. Dev Min Max
r60 201 2.373134 0.9192794 1 3
r61 140 0.0857143 0.2809469 0 1
r62a 73 2.219178 2.340742 1 15
Table 4.10: After
Variable Obs Mean Std. Dev Min Max
r60 201 2.373134 0.9192794 1 3
r61 140 / / / /
r62a 73 2.219178 2.340742 1 15

When using the summarize command in Stata, corresponding cases are automatically checked by the FDZ check script and deleted if necessary. If equivalent results are generated using other commands (e.g. table, tabstat), no automatic deletion can be carried out. In these cases, you must refrain from outputting results with insufficient numbers of cases by deleting the corresponding code from your evaluation programmes.

4.4 Percentiles / quantiles

When displaying percentiles, make sure that at least 20 observations are included in the respective percentile. A detailed output (1% percentiles) requires at least 2000 observations to guarantee the data protection requirements of the FDZ. In principle, the more detailed the analysis, the more observations must underlie the entire distribution:

  • At least 20 observations for mean values (except dummies)
  • At least 40 observations for 50% percentiles
  • At least 80 observations for 25% or 75% percentiles
  • At least 200 observations for 10% or 90% percentiles
  • At least 400 observations for 5% or 95% percentiles
  • At least 2000 observations for 1% or 99% percentiles

If you want to describe quantiles, the smallest distance between the selected percentages (also to zero and to 100) determines the number of observations needed- e.g. for the percentiles 10 – 15 – 30, the sample must be (15-10)/100 * x ≥ 20 → x ≥ 400 observations to ensure the necessary 20 observations in each quantile.

When using the summarize command in Stata, corresponding cases are automatically checked by the FDZ check script and deleted if necessary. If equivalent results are generated using other commands (e.g. table, tabstat), no automatic deletion can be carried out. In these cases, you must refrain from outputting results with insufficient numbers of cases by deleting the corresponding code from your evaluation programmes or adapting it so that percentiles with larger intervals are displayed for which sufficient numbers of cases are available.

4.5 Weighting

When using any extrapolation or weighting factors in descriptive analyses, you have to report the corresponding unweighted results. The weighted results must directly follow the corresponding unweighted results because this facilitates the data protection check and thus speeds it up - also for you. This must also be taken into account and commented on in loops. If the unweighted output is missing, the corresponding log file or, if necessary, the entire job cannot be released. Further notes can be found in FDZ template jd02_describe.do.

For weighted regressions, you have to provide the unweighted number of observations if iweights or fweights are used.

4.6 Graphs

You have to provide the underlying numbers of observations for each individual value in a graph. The total number of observations is not sufficient, rather you need to show the number of observations of each individual bar or point. You can show the numbers of observations either directly in the graphs or in the log-file by providing tables directly before or after the command for the graph. When displaying very large numbers of data points in graphs, it is not necessary to output the number of cases for each individual data point. Here it is sufficient to prove that the smallest underlying case number is at least 20. You can find an example of this in the FDZ template jd05_graphs.do.

The rule of at least 20 observations also applies to graphs. Thus, scatter plots on the individual level are prohibited, because each individual data point is based on less than 20 observations. For histograms, you need to show the number of observations for each individual bar. For kernel density plots, the total number of observations is sufficient. If the number of observations in the corresponding regression table is sufficient, coefficient or margin plots are allowed. In addition, the use of the asis option is mandatory when saving the graphs in gph format to ensure that the graphs are saved in their current form and that no further editing is possible. Further instructions can be found in the FDZ template jd05_graphs.do.

4.7 LaTeX output

Stata result files must have the ending “.log” or “.txt”. Results written in separate files outside the log files created by Stata cannot be released. You can include the results of these separate files in the log file with the command “type path”. Make sure that the output is clearly arranged and easy to read (e.g. through fixed column widths). If this is not possible in exceptional cases, the results must be displayed again in an easily checkable form directly before the included text. An example can be found in the FDZ template jd03_analyses.do.

4.8 Aggregated data

4.8.1 Importing aggregated data into the project

External data on an aggregated level (e.g. unemployment rates by district) may be merged to the FDZ data if they comply with the data protection guidelines. For more information on the requirements and the procedure, please refer to our manuals (section 2.2.2).

4.8.2 Documentation of aggregation steps within the project

If you aggregate the original microdata for further analysis (e.g. to the level of industries, regions, etc.), please make sure this is properly documented in the log-file (captions; explanatory notes). All log-files created after such an aggregation step (and using the aggregated data set as the main data source) must contain the following information in the top header:

  • the level of aggregation,
  • the programme step (i.e. do-file), in which the aggregation takes place (please also mark this step in the master file).

If this information is missing or hard to find, the results cannot be checked and released.

If you want to display those aggregated variables later, note that the rules for the minimum number of observations still hold (see chapter 4.1 to 4.3). Therefore, the following requirements apply:

  • An additional variable must be created for each aggregated variable to be displayed, which contains the number of the corresponding non-missing observations in each cell.
  • If the aggregate variable is a quota, share or rate, tabular representations must also show the number of non-missing values in the individual subgroups used to construct the outcome (e.g. the number of men and women in federal state X alongside the women’s share in federal state X).

Examples on how to generate these variables correctly can be found in FDZ template jd02_describe.do.The number of aggregated tables should be kept to a minimum to comply with the principle of data parsimony1.

4.8.3 Export/transfer of user-generated aggregate data sets

If you wish to receive a more extensive aggregated data set as a .dta data file because it cannot be properly displayed and reviewed as a table in a log-file (or if you would like to transfer that data set to another project directory), please discuss the procedure with us in advance, ideally already when submitting the application. It must be clarified what the purpose of the aggregation is, at which level it is aggregated and how the included variables were generated (sum, mean, etc.). Explain why the export/transfer is necessary.

The generation of the aggregated data can be prepared with the test data and should be done during on-site use. The programmes must then be started again with JoSuA. Please use the Presentation/Publication Mode for this and note in the comment field that you want to receive or transfer a generated data set and the corresponding file name. Always save this data set in the subdirectory data.

Since checking aggregated data sets is very time-consuming, the following rules must be implemented when creating them:

  • The entire data preparation for this aggregate data set, from the original data to the export data set, should be performed in one do-file/log-file. Please omit any unnecessary steps in this do-file. An example can be found in the FDZ template jd06_export_aggregate.do.
  • The rules outlined in the previous chapters concerning documentation and reporting of number of observations apply equally.
  • Use informative variable names, variable labels and value labels for the export variables.
  • Avoid stacking aggregations (e.g. state + labour market region + districts), at least if the finest level of aggregation has cells that need censoring.
  • Do not add variables for additional subgroups using the ‘wide’ format. Keep the format ‘long’.
  • The variables defining the aggregation level have to be named “lvl_1_*, lvl_2_*, …”, e.g. an aggregation to year and federal state should lead to grouping variables lvl_1_year and lvl_2_fstate.
  • Each aggregated outcome variable has to be followed by a variable containing the respective number of observations used for constructing it (if it is not a count variable itself). This variable has to be named the same way but starting with an N_.
  • Try to keep the aggregated data set as small as possible. In particular, values that can be calculated from other information in the aggregated data set (e.g. totals or ratios) should be generated after export or transfer to simplify disclosure control.

In addition, for the export/transfer of an aggregated data set, you must ensure that each cell is based on at least 20 observations. If this is not the case, you have the following options:

  • Group categories (preferred by the FDZ). For example, if you want to obtain aggregates at district level, group together districts that have too few cases.
  • Censor cells that have too few cases yourself (not preferred by the FDZ). Proceed as follows:
    • If the number of observations in some cells of the data is < 20, set the count variable and all other outcomes for that cell to missing category “.p”.
    • Keep the number of cells with insufficient observations at minimum. Keep in mind that our staff might have to perform secondary suppression (see chapter 4.2).
    • Add an explanation why further aggregation/grouping is not justifiable from a project standpoint.
    • Please note that we reserve the right to decline releasing data for export or transfer if it contains low cell counts that are not appropriately justified and documented, or we decide that the time needed for appropriate disclosure review would be excessive.

For each project, such a data set can be forwarded only once. This also applies if you want to have an aggregate data set transferred to a different FDZ project folder in order to merge it with other data sets. The export of the data set is done by email. Transfers to other project directories are carried out by FDZ staff members.

4.9 Regression output

The total number of observations used in regressions must be at least 20. Results of multivariate analyses based on fewer than 20 observations can generally not be released and should therefore be removed from the programmes.

However, with the help of regressions, small numbers of cases can also be determined indirectly under certain circumstances. Therefore, in these cases it is not sufficient to consider only the total number of observations. The following regressions are problematic:

  • Regressions whose coefficients result in the unconditional mean, such as in a regression with only one dummy or a group of dummies, e.g. for the federal states.
  • Regressions with two or more variables in fully interacted models (all possible combinations of variable values).

Here it must be ensured that the individual values of the dichotomous or categorical variables are based on at least 20 observations. This is to be shown by frequency tables of all variables after the regression output using if e(sample). Each additional variable in the model no longer allows for conclusions. Further instructions can be found in the FDZ template jd03_analyses.do.

4.10 Event data analysis

When using the sts list command in event data analysis to identify the Kaplan-Meier estimator results, the value at the beginning of the table (population at risk) is relevant for the data protection check. In addition, when groups are shown in graphs, the sts list command must also be used for each group to provide the underlying number of observations.

When results from survival analyses are displayed, only the initial value must be >= 20, not every single step.

4.11 Sequence pattern analysis

Graphical representations of sequence pattern analyses must be based on a sufficient number of observations. There are two ways to do this. First, you can show sequence patterns per unit of time as proportions of states (e.g. in month 1, 30% of people are in unemployment and 70% in employment). Second, you can depict aggregated sequence patterns for groups if each group contains at least 3 observations. Graphical representation of sequence patterns for less than 3 persons or establishments or enterprises is prohibited.


  1. It should be noted here that the automated check script also deletes or censors aggregated outputs using certain Stata commands such as summarize or tabulate, for example at the district level, despite a sufficient number of persons or establishments or enterprises within the districts. In such cases, we would ask you to contact the FDZ in order to avoid such a deletion of results through adapted programming.↩︎