Architecture Concepts of Databases

This is a guide on architecture concepts of databases. 

Check out Audible on Amazon and listen to the newest books!

Introduction to Databases

A database can be understood by comparing it to a physical filing cabinet. A filing cabinet contains some type of organizational system that allows a person to enter information in a way that makes sense to someone who will later need to retrieve that data. If every single file were simply piled into any given drawer without any organization, locating specific information when it is needed would be extremely difficult.

To solve this problem, a system is created in which drawers are labeled and folders and subfolders are used to organize the data. Databases function in a similar manner. They can be processed to deliver meaningful information to end users, which in most cases means selectively examining data rather than examining all of it. A user may only be interested in certain customers, certain orders, or certain employees that meet particular criteria. In a simple example, a user might only want to examine employees who have the Platinum benefit package.

Those specific records can be extracted so that whatever other type of analysis is needed can be performed. Data can also be manipulated in various ways when a change needs to be made. This logical structure allows data to be entered and retrieved effectively and efficiently.

Data within a database is organized first by tables, which are sometimes referred to as relations. A table is simply a collection of records as a whole. Tables are used to store all of the records in that particular table, and they are also structured and organized.

Tables or relations also identify where data is stored in a file. They break down information into what are known as fields and records. Fields are also called a data item or element. The term tuple may also sometimes be encountered. A field is essentially a vertical column. In an example table, the fields might be Employee Code, First Name, Last Name, and Benefit Package. Any individual value in that field is the smallest unit of data in the table. For example, Employee Code 0648 indicates a single employee, and that is the smallest unit of data representing the code for that particular employee.

A record is read across the horizontal, and it puts each individual field together. A record, also called a tuple, is a set of logically related fields. For example, Employee Code 0648, Keith Daniels, and the Gold Benefit Package together form a complete unit, which is one single record in the table. Records generally have fixed data types as well as fixed values.

A data type is the kind of information being entered. In the example, the Employee Code is numeric while all of the other fields are text based. These are examples of the types of data. Most tables want to know what kind of information is being entered, such as numeric, text based, image based, date, or time based. These are just some examples.

The value itself is typically fixed, but this does not mean that it cannot change. Keith Daniels, for example, may receive a better benefit package after a few more years of service. Therefore, the value certainly can be changed. However, once it is entered, it is usually fairly static. It does not mean that a smaller environment cannot have one; it is just not as common.

Databases usually contain large amounts of static data, which means that the information is no longer going to change. Even if it might have changed at some point, the data warehouse is usually referred to as an archive. What is in there does not change anymore. Therefore, there are hardly any writes to a data warehouse; it is mostly reads. That data is no longer required for day-to-day operations.

In a typical data warehouse graphic, several data sources are shown. These represent the still ongoing day-to-day operations once that information is able to be archived. As an example, one might think about sales. The day-to-day sales might be the current sales for this year, but sales for last year, the year before, or the year before that are not going to change anymore.

However, those sales can be analyzed to determine various types of information. What ends up happening is that the data is generally loaded from the original data sources by first performing what are known as ETL operations, which stands for Extraction, Transformation, and Loading. Extraction pulls the data out of the current data sources.

Transformation involves possibly reformatting some of the information, particularly in terms of data types because they may not match, so conversions and similar operations can be performed. Loading is simply getting the data into the data warehouse. Once the appropriate information is in the data warehouse, users can start analyzing that information, such as what causes a product to sell more effectively.

A new record can be inserted, an old record can be deleted, or an existing record can be queried, usually fairly quickly. Relatively few records are returned when a query is run, as opposed to hundreds of thousands of records; a user might just get a few hundred.

These databases are highly normalized, which means very organized. A lot of thought went into the structure of the tables themselves. This might be something like orders. For example, if somebody has an online store, the people generating the orders are simply entering records into this table as they make their orders. That is what is referred to as online transaction processing.

This rapidly executes very complex queries. Such systems can answer questions involving planning, problem solving, and decision support. They are widely used in business reporting, budgeting and forecasting, financial analysis, and sales and marketing management. This is what was referred to earlier when discussing the attempt to ascertain useful information that is not readily available just by looking at the raw data.

All of the records have to be examined in a complex fashion so that when one thing happens, it results in another thing happening, and forecasts, analysis reports, and similar outputs can be created to decide how to move forward. That is an example of decision support. These are the primary types of databases that are likely to be encountered in a database environment.

A DBMS stands for database management system, and it is ultimately how users access the database. It allows users to add, delete, update, and retrieve data. To clarify quickly, it is for starters a software application. Any kind of DBMS is simply a piece of software that is installed on a system. End users generally do not directly interact with the DBMS.

As far as the DBMS interface itself goes, that is generally only accessed by database administrators, database developers, and possibly application developers themselves. Therefore, it is usually not a front-end user tool. It is, of course, used by the database to manage the data itself and also sits between the database and the operating system on which it is installed.

It is this overall intermediary, not only between the users and their applications, but also at the lower level, the hardware and the operating system. It basically controls everything with respect to the database and how it operates. It is connected to the applications and the users that are using it. It is also connected to the operating system and the hardware that the database relies on.

The DBMS helps define what are known as schemas and subschemas in a database. Schemas are the blueprints, the way that the overall database is designed, and simply the things that can be done with that database. It is a fairly low-level type of object, but it really comes down to how the database is simply created and defined. It allows relationships to be established between different data elements.

This is how information in one table can be connected to information in another table through what are called relationships. It absolutely allows data to be entered into the database and then manipulated and processed, and it also handles security, limiting access to only authorized users.

Information can then be retrieved by querying the database. This requires a language known as SQL, or Structured Query Language. This is how information is pulled out of the database when records need to be retrieved. Some of the DBMS tools that are available are numerous.

The end user, as mentioned, typically goes through some kind of application program, a front-end interface. That interface will then typically access any number of DBMS tools that will submit the query language to the DBMS software and ultimately down through to the database kernel. Query languages include things like SQL, but that is not the only one. There are a number of implementations, flavors, or dialects of query languages that are available. They pass the query statement down through to the DBMS software.

Imagine a user wanting to retrieve a record. That user might need to supply some kind of criteria. For example, a salesperson might want to see the customers in that salesperson's sales territory only. The salesperson would enter something like their own identity and then the sales territory they represent.

That is referred to as criteria. It has to get packaged up into a query language statement that is then passed through to the DBMS software. The DBMS software parses that statement and retrieves the results. It also has to operate right down to the kernel level, or what is known as the database engine, to ensure that it retrieves only the appropriate records that meet the criteria and also that it does not interfere with any other users who might be accessing the same data or possibly updating some of that data while the retrieval is taking place. Those are what are known as kernel-level operations, the ability to prevent users from interfering with each other. It is just one example of that.

A data dictionary might be included, which allows terms in the statements to be literally defined to make sure that valid statements are being issued to the engine. Ultimately, the hardware and the operating system are involved because things like storage and how much memory is allocated to the database itself have to be controlled. All of that can be configured within the DBMS. However, that is typically not something that end users are concerned with; it is more so the administrators and the developers.

The DBMS advantages include consistency. If a particular customer is retrieved by their ID, that customer will always be seen. This does not mean the customer might not be deleted, but it is always going to be a consistent record. Some information might change; for example, people move, so their address changes, but the customer as a whole is still the same customer, and anywhere else that customer is referenced in the database always refers to that same customer. Therefore, the records remain consistent.

Limited redundancy means that more information than is necessary never has to be entered. Once customer information has been entered, for example, if an order then needs to be created for that customer, all of the customer information does not have to be redundantly entered. Only a specific and usually single value needs to be referenced to identify who that customer is, such as an ID.

When the order is created, it does not say it belongs to this person with this first name, that last name, that address, and that phone number; that information has already been stored in the customer table. The only thing that needs to be referenced is something like the customer's identity, the single unique value that specifies that particular record. That is limited redundant information.

The DBMS maintains security over the data to ensure that only the appropriate people are accessing it. At the same time, it facilitates data sharing and concurrency so that multiple people can be in the database at the same time and can still access the same records at the same time.

For example, if records are simply being selected, multiple users can select the same information at the same time without causing any problems. That allows for data sharing capability. Independence from application programs is another advantage; usually the DBMS does not care what kind of front-end application is being used. It also facilitates recovery and backup.

A backup of the database absolutely needs to be made in case it fails. Finally, efficient accessibility means being able to very effectively and efficiently enter data and retrieve data very easily in a manner that makes sense to users.

Database environments typically have a number of different types of users for accessing the database for a variety of purposes. Depending on the skill level of those users and what they ultimately need the database for, different categories can be determined for those users.

They can be grouped into DBAs, which are database administrators, database designers, end users, and application developers. Any one person is not necessarily relegated to just one of those categories. They might have extensive experience with all of them. As mentioned, it depends on the skill level of the users. Each of these will be examined, starting with the database administrator.

The database administrator is responsible for controlling and managing the overall database and the database management system. They are responsible for overseeing software and hardware resource requirements, ensuring that there is sufficient processing and memory and disk space, for example, for the database to maintain its performance, and also ensuring that the appropriate software is installed to meet the demands of the database and, of course, the end users who are accessing it.

The main day-to-day responsibilities include regularly monitoring performance and resolving any issues if things start to slow down. They authorize users to access the database, thereby controlling security, including which users have access to which tables. They determine the contents of the database and how data is stored, looking at the structural elements and determining where a file should be placed versus where a transaction log file should be placed.

They protect the database from hackers and viruses, which again relates to security. They ensure the data is consistent and valid, so data integrity needs to be maintained by the administrators. They ensure there is sufficient space available; some databases can be exceptionally large and require very advanced storage solutions. They also back up and restore the database so that protection against failure is ensured.

Database designers are responsible for the overall design of the database, the architecture. They often liaise between users, management, and other business personnel to determine the requirements for the database.

Not only do they need to be good designers, but they also have to understand the business processes that are required for the organization. They then design and create the storage structures, both physical and logical, including the actual hard disk storage requirements and how the tables interact with the files that store them. Finally, they design and create the tables in most cases, the constraints to control the data, and the relationships between the tables.

End users are the ones who are in their day-to-day querying the database to retrieve data, as well as adding new records, updating existing records, and removing old records. All of that is basic data manipulation. End users can be classified into four different types: sophisticated, parametric, casual, and standalone. There may be some flexibility in those terms.

A sophisticated user is someone who is thoroughly familiar with the database and how to access it using a language such as Structured Query Language. They probably have some back-end or possibly even administrative experience. A parametric user is generally not all that aware of the database itself. They may be very familiar with the front-end application they are using, but it is using canned transactions.

They are just clicking some buttons to submit queries, or they might have to type in some criteria or search for a customer. However, all of the transactions they are invoking are done just by clicking buttons on the interface. A casual user is someone who infrequently accesses the database. This is not so much a skill level as it is a role, such as a temporary worker or a contractor hired temporarily just to do some data input. A standalone user is generally someone who is using a personal database for their own requirements, not necessarily a server-class database.

Application developers are responsible for designing and creating the database application, which is the software that the end users are using. They need to have a pretty good understanding of the database structure for the application and how the storage works because they need to make sure that their application is accessing the data.

They provide information on the requirements for the database structure to support the application. There needs to be a very solid understanding of the database environment so that the application is gaining access to the data appropriately. Finally, they tune and resolve issues with the application itself and its access to the database. They often work with the DBA and possibly the designer in those cases to ensure that everything is working as well as possible.

They may be seen working together at the very start of a project or really any time afterward. Most of the time, they will all be working together during the implementation of a database. Once it has been implemented, from that point on it is typically the DBA that looks after everything and, of course, the end users who use it. The application developers and the designers are usually more involved at the point of implementation and not day-to-day operations, but those are all the different types of users that may be encountered in a database environment.

In knowing that a database requires a fair amount of planning and architectural considerations before it is implemented, what then makes a good database? There are poor databases out there. There are a lot of things that need to be in place to qualify a database as good. Some of the things to look for are certainly critical for the database to be available and for the information to be retrieved in a very consistent and accurate fashion.

Available basically means that users are not trying to gain access to the data and simply unable to connect because the server went down or something along those lines. Maybe the server is fine, but a connection simply cannot be established due to network errors or anything along those lines. Any kind of downtime can be critical for some environments.

They need to have that database available to them, and the information that is coming back needs to be consistent and accurate at all times. It also, of course, needs to meet the needs of the users. That can be a bit of a tricky process to nail down, especially right at the beginning, because things will ultimately change a little bit, and even business processes may evolve. That can be a bit of an ongoing process, but the best possible effort certainly needs to be made at the time of design and implementation to meet all the requirements of the users.

A good database is accurate in terms of the data. It performs at acceptable levels, which is a little bit arbitrary as well, but users definitely should not always have to be explaining to people on the phone that the system is running slow today. That is just not acceptable.

A good database efficiently accesses and uses its storage. Databases can get very large, but the storage should be very efficient as well. Databases typically grow in increments. Those increments should be efficient and should not be grabbing huge amounts of storage for no reason and wasting a lot of space. The data itself should only be stored in blocks or chunks that are just large enough to store that data and, again, not wasting a lot of space.

For example, if a character-based field only requires two or three characters, the data type should not be 50 or 60 characters because that would waste a tremendous amount of space. The database certainly needs to be secure. It needs to provide access to the data when required, so users are not sitting there waiting and waiting to get access. It also has minimal redundant data.

Once a customer record has been entered, for example, the information about that customer should really never have to be entered again, with the exception of having to refer back to their primary key value, which uniquely identifies them. Once that value is available, everything else can be retrieved about the customer. Their name, address, or phone number should never have to be entered redundantly. It has already been entered in the customer table.

As far as data accuracy goes, data should be accurate and free from error. This can be a little bit tricky because users may inadvertently enter inaccurate data if the correct measures are not in place. A simple typo will happen from time to time. One example is date fields in particular. The example 6/7/15 could be June 7th or July 6th. Measures really need to be in place so that the environment and the users know the format being used.

This can cause problems, particularly when querying data with people supplying the incorrect format. The records expected to be returned simply will not be returned. To ensure data accuracy, a good database should include data integrity constraints preventing incorrect values from being entered. Things like pick lists on a front-end application really help to eliminate those issues, particularly where dates are concerned or maybe regions.

A list can be provided from which the user picks rather than allowing the user to type it in. Foreign key constraints for referential integrity always refer back to a primary key value to uniquely identify a record. Additional constraints on columns, such as default values and unique constraints, prevent duplicates and enforce formats for things like dates.

Performance is critical in a good database to ensure that data is retrieved in an acceptable amount of time. Adding and updating data should be performed in a timely fashion to ensure data consistency and accuracy. If data is not managed in an acceptable time, issues may arise with users receiving outdated data or having to wait long periods for it to be returned.

If it takes forever to pull a record back, people are given more time to enter other records that may have otherwise been retrieved if they had been entered a little more quickly or retrieved a little more quickly. To exaggerate, if a user wants to see a bunch of customers in their territory and pulls them all back, but it takes a very long time for that to be accomplished, somebody may have already entered new records into that territory and they will not come back.

Alternatively, it takes forever to enter them and they are retrieved too quickly before the other records are entered, so that can be tricky as well. Some of the factors that affect performance include network speed, the design of the database, the design of the application, and, of course, the hardware supporting it.

In terms of storage, adequate storage absolutely needs to be available because databases can become very large. Determining the amount of storage required depends on the amount of data the database is projected to store, which can be difficult to know sometimes, as well as the software itself and the operating system requirements. To avoid limitations, old data can be archived out or more storage can simply be added, which is usually fine in most cases.

Insecure databases can cause data to become corrupt or allow sensitive data to be stolen. A good database only allows authorized users to access it. Each user should always have their own unique ID and password with specific privileges to the data they are able to access. Database access needs to be monitored to ensure that no unauthorized users are getting in and that those who are authorized are only gaining access to the data they should be able to view.

Availability refers to being able to access the database at all times. Anytime a connection is needed, it is possible. To ensure the database is available, it should be well designed, and downtime should be scheduled outside of peak business hours with minimal downtime or even zero downtime by using a variety of high availability techniques, which usually include multiple redundant copies of the database.

If one fails, there is another copy ready to assume services. Advanced applications that query and update the database are used, and those applications do not require the database to be locked or shut down to perform those updates. This typically is not something like inserting a single record, but rather importing thousands of records, which can take a very long time.

Redundancy is another consideration. A good database ensures that redundancy is minimized. This can make data very efficient and basically help ensure that inconsistencies and inaccurate data are avoided by minimizing the amount of redundancy.

Redundant information can simply take up space that could otherwise be allocated to good data. Minimizing redundancy helps free up space, avoid inaccuracies and inconsistencies, and, of course, save time in processing the data. Once a customer is entered one time, for example, that customer should never have to be entered again. The single unique value that identifies them is entered when they need to be referred to, and basically that is it. These are a lot of the characteristics of a good database.

Database management system architecture refers to a set of processes, rules, and specifications that describe the nature of the data itself, how the flow of data is controlled, how the data is used by applications, how the data objects are stored, and how the data is integrated.

All of these put together help ensure the reliability, integrity, performance, and scalability of the database. Different types of businesses, of course, use different types of architectures depending on their requirements. Some of the commonly used architectures include centralized, client/server, N-tier, where N basically means any given number, distributed, and parallel. Each of these will be examined, starting with centralized.

The centralized architecture is certainly the oldest form of architecture and was commonly used in mainframe environments. These are not seen a whole lot anymore. In the mainframe world, a dumb terminal was used to connect the user directly to the mainframe computer itself.

The mainframe is a very large server-type system that stores and processes data and applications. The terminals themselves did not have any kind of processing power. There was no operating system on the terminals. They only connected the user to the mainframe and displayed the data. All of the processing was still happening on the back-end mainframe, hence the term dumb terminal. They really did not do anything other than display the information.

This was something typically seen when simplicity and cost were the driving factors rather than scalability and flexibility. It was pretty easy to implement additional dumb terminals, for example, but significant changes to scaling the mainframe really were not feasible.

It was not particularly flexible. Once it was constructed, it was pretty much as is. It is not that changes could not be made, but it is much easier these days for a database to grow in scale than it was using this type of solution. It was fairly easy to deploy and maintain, though, and all of the processing was completed on the server. There was virtually nothing that had to be done on the front-end side. There were no applications to install, and there was no maintenance involved other than physically replacing a terminal, for example.

There were no updates and really no day-to-day administration of the front end. It was all done on the back end. It was certainly a little easier in that respect in terms of deployment and maintenance.

It is certainly, however, as mentioned, becoming far less common today. There is limited operational capacity for handling complex, scalable, and flexible processes. With hardware becoming less expensive, businesses can afford client PCs, replacing dumb terminals and workstations, and even that in and of itself is something that has been around for a very long time.

The centralized architecture is certainly far less common these days. It is not extinct by any means, but it is probably the least common at this point in time. It still had its advantages and sufficed for a very long time. However, the fact is that hardware is a lot cheaper and servers are a lot easier to implement than they were with respect to pulling a mainframe into an environment. Mainframes were typically very large and took up entire rooms worth of space. There are really not a lot of mainframes left, but they certainly served their purpose in their day.

The client-server DBMS was designed to improve usability, interoperability, scalability, and flexibility. These days, it is probably still one of the most common implementations of database architecture. It contains three primary components despite the fact that two-tier is often discussed.

There is a third component: the client itself, the network interface, and the server make up the three components, but there are still two levels of processing. The client is used by end users to access the database. The client interacts with the server to process queries and retrieve data and contains the client components for the DBMS.

This is both the first component and the first tier because what is now present is an application that can be installed on the client and is used to access the database, which becomes the tier. The database itself residing on the server is also a tier, and the network interface simply connects the two. Therefore, there are three components, but just two tiers.

The network interface uses the various protocols required to connect the client to the server. That can be both a networking protocol just to transport the data and a database protocol to submit the appropriate statements to the database.

The server, of course, houses the database management system and is dedicated to managing the enterprise data, monitoring the clients, and allocating the workload as well, but it basically has no concern for the application that is being used. It simply stores the data and manages access to it, but really does not care about the applications on the client. Again, those are the two separate tiers.

As far as advantages go, the client-server architecture improves the performance of the database management system. It shares the network load among multiple clients, which helps reduce server-side processing because there is an application on the client that contains an operating system. In other words, it is not a dumb terminal. It does a lot of the processing itself, which alleviates the workload on the server. It provides a user-friendly interface for end users, which again is the application they are using. It supports development of complex applications because the application can run entirely on the client; the server is not burdened with the complexity of the client. The client can run the application just fine. It supports clients that run different operating systems because the server does not really care what is accessing it. It also reduces the overall cost of the database management itself.

There are still some disadvantages, and this does depend on the implementation. Things such as a single server can be a problem. If there is only one database server, then failure of that server will prevent clients from connecting to the database. If it goes down, everybody is down.

This can be mitigated, for example, with things such as a secondary server or high availability solutions, wherein if the server goes down, there is another one, a redundant copy ready to assume the server role. This is also being mitigated these days by implementing things like remote applications or web-based interfaces that do not have to be deployed; a browser is simply used. Those disadvantages are still there, but there are certainly ways to mitigate those concerns.

The most common implementation of N-tier DBMS is using a web-based application or a web page to gain access to the database. Hence, the application layer shown in a typical graphic consists of web servers.

They have nothing to do with the database itself, nor do they really host any kind of software; they simply host websites. The clients access the database through that website. Therefore, the application layer resides in between. There is a client layer, one or possibly more application layers, which are referred to as middleware in a lot of cases because they reside in between, and finally, a database server layer.

At the application layer, the transfer of data between the client and the server is still seen. It still controls user access to the database servers. It shares the tasks of clients and servers, connects existing systems to new systems, and enables clients using different network protocols to access the database.

Some of the advantages here include flexibility to mix and manage various programs, applications, and interfaces in different tiers. In fact, this can be intermixed with a client-server architecture because there might still be internal clients who are accessing the database server fairly directly in a client-server model, while external users are accessing it through the web interface, through the application layer. Therefore, the architectures themselves can be mixed and matched.

Scalability is another advantage. Application servers can be deployed on several systems. Three separate web servers can balance the load of all of the requests. If it is too much for one server, another server can be added, basically creating what is called a server farm, or a cluster, which is another term.

They can all distribute the workload across each other, so more and more servers can be added as demand increases. Network intelligence is another advantage. Division of tasks is more intelligent and faster, which enhances processing speed and frees up memory space. Compatibility is also an advantage. Applications themselves are not installed on the clients, so they are managed from the application layer only. In essence, the clients are accessing the database through a browser.

Almost every system has a browser, so no software has to be deployed to the clients other than a browser if they do not happen to have one, which is pretty rare. In terms of managing it, all that needs to be done, for example, if an update is wanted, is to update the website at the application layer.

Changes are simply made to the web servers, and every client sees that same website. The change has been deployed to every client. Therefore, there is far greater compatibility and far greater manageability by implementing changes at the application layer only.

In a distributed DBMS, clients can access the data from any known location without necessarily even knowing where it is coming from, so it is entirely transparent to them. This comes back to the ability to operate independently if desired.

In the case wherein the database becomes inaccessible at a user's own site, the user could in fact access the data from one of the other sites because they are still connected through a network. With this, two types are typically seen. Homogeneous means each location runs the same software, same database management system, same clients, same database software, and so on.

The opposite of this is heterogeneous, wherein each location could in fact run different software, different DBMS, and different clients. It really depends on the environment. In most cases, the database itself is still the same. It is just how the database is accessed and managed that might vary.

There are some advantages to this. Multiple levels of transparency for the data exist. Users can access their local copy first, but should they lose access to that, they can get it from really anywhere else. This increases reliability as well. The chances of all four sites being down at the same time are pretty low.

This provides for local autonomy because a site can operate entirely independently of the other sites. If site one is operating and sites two, three, and four all go down, site one still has its copy, so it can still operate. There can be security for sensitive data, so if it needs to be isolated to a certain degree by site, that can be done. Better performance is another advantage because the local copy will always be accessed first over a LAN as opposed to using WAN or Internet links to access a copy in another site.

Ongoing operational facility is another advantage because if connectivity is lost at a site, other sites can still be accessed, so overall operations can usually continue. Easier expansion is also an advantage because another site can be added really at any time.

There are some disadvantages to this. It can be complex, so it can make it difficult to manage. A little extra effort and cost might be required. There could be additional software to install and deploy at each site. That will in turn result in possibly additional maintenance that is required and increased cost for the additional hardware in each site and, of course, possibly increased resources as well.

The last type is the parallel architecture, which is included with the distributed because it is based on the principles of distributed but shares the resources of the servers. The memory and the disks, for example, can all be shared so that what is known as a pool of resources is created.

This way all of the memory across all of the servers is pooled up together, but then divided generally equally, but not necessarily. One particular system may require more than what it actually has, so it can draw more from the pool. If there is a system that is being fairly underutilized, its resources can be shared out with any system that is being overutilized.

The main architectures within this are shared-memory, shared-disk, or shared-nothing, and this simply refers to the resources that are being shared. Memory, as mentioned, just pools all the memory together and then divides it up as equally as possible. The same with the disk: it pulls up all the disk space and then allocates that as best as possible.

Shared-nothing typically means reverting back to distributed. It is usually something that is used maybe just temporarily, turning off the feature to make some adjustments or similar operations. Any particular system can be isolated at any time, and that is the shared-nothing approach. As mentioned, it is usually something that is temporary, not something that would be implemented in a permanent fashion.

These are all the different architectures, and there are certainly different implementations for each one of them. It really comes down to whatever is felt to best suit the needs of the organization.

There are several different database models, but one thing they all have in common is that they attempt to apply some kind of structured representation of the data to the information that is stored. This gives a means for storing and organizing the data elements, particularly when records are being entered or, perhaps even more importantly, when records are being retrieved.

If there is no organizational method, if there is no structure, it would be particularly difficult to locate a specific record or even a collection of records. There are seven models: flat file, hierarchical, XML, network, relational, object-oriented, and object-relational. Flat, hierarchical, and XML will be examined first, and the later ones will be examined subsequently.

The flat-file database model uses a file that is a simple text format. A sample cut out of Notepad contains three lines, each separated by a semicolon. Line one is Employee, colon, 657, colon, Roy Anderson, colon, 906, North Main Street, semicolon. Line two is Employee, colon, 780, colon, Calvin Boyd, colon, 620, Mission Street, semicolon. Line three is Employee, colon, 804, colon, Cathy Patterson, colon, 115 Trinity Avenue, semicolon.

With this, the data really is unstructured, at least in terms of an application or a software level. There is nothing in here that is telling what is being looked at. There are no field headers, for example, but at a human level, it can be inferred that this is employee information, some kind of ID or code, the person's name, and the person's address. Even though there is no real structure to this, the data is still stored in some kind of a field, and in this case, separated by a delimiter.

The delimiter is simply a character that indicates the end of that data. At the end of the record, a different delimiter is seen, which indicates the end of the record. The colon in this case is the field separator, and the semicolon is the record separator. Certain applications, if an import of this information into a more structured database were to be done, can actually pick up on those and determine that is the end of this field, then move to the next field.

Once the semicolon is seen, that is the end of this record, so move to the next record. Even though there is not much structure to this, it can at least be looked at and known to a certain degree what is being looked at. However, it would be fairly difficult, particularly if this was a very large file, to find specific values. This is useful for storing simple data. Once it starts to get a little more complex, flat-file databases usually do not tend to work all that well.

In the hierarchical database, the data is stored in the form of a tree structure with parent and child files. They are connected using either one-to-one or one-to-many. This simply means that one entity might be connected to only one other entity in the one-to-one relationship, or that one entity could be connected to many other entities in the one-to-many relationship.

It does not really matter which one is used; either one can be implemented. With this, the data is organized into a specific hierarchy that has a flow to it. For example, there is a project. That project is being handled by two different teams.

Each team handles the same entities of developer, tester, and quality control, but some kind of separating criteria would be imagined as to how somebody becomes a developer of team 1 versus developer of team 2. Maybe it is geographical, maybe it is by department. It does not really matter in this case, but there is something that separates